39% of earnings calls trip the complexity meter - and they raise guidance less often
This study examines earnings calls scoring 6 or higher on a 0-9 language complexity meter. Of 165,182 calls from 1990-2026, 64,523 (39.06%, 95% CI 38.83%-39.30%) qualify. High-complexity calls show +0.34 higher evasion and +0.32 higher stress, but -0.14 lower confidence and -0.11 lower candor than typical calls. Only 16.22% of them raised guidance versus 21.05% of other calls. Their next-day returns are weaker: a median of -0.0894 versus -0.0716, with 36.98% beating the market versus 39.47%. Complexity peaked at 44.62% of calls in 2015 and fell to 32.77% by 2025.
- 64,523 of 165,182 calls (39.06%) scored 6 or higher on the 0-9 complexity meter.
- High-complexity calls score +0.34 higher on evasion and +0.32 higher on stress, but -0.14 lower on confidence than other calls.
- 16.22% of high-complexity calls raised guidance, versus 21.05% of all other calls.
- Calls following high-complexity language had a median next-day return of -0.0894 versus -0.0716 for all other calls, with 36.98% beating the market versus 39.47%.
1Introduction
Anyone who reads earnings transcripts knows the feeling of a call that sounds elaborate but says little. Complexity in management language is easy to notice and hard to quantify, which is why a simple 0-9 complexity meter is useful: it turns that impression into a number that can be measured across decades of calls. This study asks what distinguishes the 39.06% of calls that score 6 or higher - how their candor, evasion, stress, and confidence differ, how often their companies raise guidance, and how their next-day returns compare with the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 complexity meter (n = 64,523; 39.1% of the reference set, 95% Wilson interval 38.8%–39.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
High-complexity calls read differently: evasion is +0.34 higher and stress +0.32 higher, while confidence (-0.14), candor (-0.11), and specificity (-0.08) all sit below the corpus norm. The guidance pattern points the same direction - 16.22% raised guidance versus 21.05% elsewhere, and 12.70% lowered it versus 11.56%. Returns echo the tone gap: median next-day return -0.0894 versus -0.0716, and a 36.98% market-beating rate versus 39.47%. One phrase is overrepresented by 1.58x: 'Scale-Dependent Advantage Claims.' The trend is also striking - complexity fell from 44.62% of calls in 2015 to 32.77% in 2025, with a sharp drop around 2020 (35.91%).
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.75 | 6.86 | -0.11 |
| Evasion | 3.04 | 2.70 | +0.34 |
| Specificity | 7.47 | 7.56 | -0.08 |
| Stress | 2.75 | 2.43 | +0.32 |
| Promotion | 5.10 | 5.05 | +0.05 |
| Confidence | 7.07 | 7.21 | -0.14 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 16.2% | 21.1% |
| Maintained | 49.7% | 48.8% |
| Lowered | 12.7% | 11.6% |
| Withdrawn | 2.3% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 1.58× | 17.5% | 11.1% |
| The Question Left Hanging | 1.31× | 62.8% | 48.0% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -8.9% | -7.2% |
| Interquartile range | -27.9% to +10.1% | — |
| Share beating SPY | 37.0% (95% CI 36%–38%) | 39.5% |
| Observations | 8,659 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| HCA | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FLG | Q2 2025 | 2025-07-25 | B |
| TNET | Q2 2025 | 2025-07-25 | C+ |
| DBOEY | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude that high-complexity calls co-occur with more evasive language, less raised guidance, and slightly weaker next-day returns. That is a pattern of association, not proof that complexity causes anything - complexity may simply travel with harder quarters, tougher questions, or stressed sectors. The return gap (-0.0894 vs -0.0716 median) is descriptive, not a signal to trade on. The right takeaway is that language complexity is a measurable lens on management posture, one worth tracking alongside fundamentals rather than instead of them.
5Limitations
Complexity, candor, and stress scores come from AI-read transcript fields and are noisy by nature. The returns comparison covers 8,659 calls within a base of 22,449, skewed toward liquid names, so medians and beat rates may not generalize. Our own forward tests falsified directional prediction from these features - the return gaps shown here are historical descriptions only. Finally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of language-based scoring against realized outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.