Only 11.3% of earnings calls score high on uncertainty - and just 4.6% raise guidance
This study examines 18,704 earnings calls (11.3% of a 165,182-call corpus spanning 1990-2026) that scored 6 or higher on a 0-9 uncertainty meter. High-uncertainty calls show sharply different behavioral profiles: stress runs 3.66 versus 2.43 baseline, evasion 3.31 versus 2.70, while confidence drops to 5.96 from 7.21 and promotion language falls to 4.15 from 5.05. Guidance behavior diverges markedly: only 4.6% of these calls raise guidance against 21.1% baseline, while 32.6% lower guidance versus 11.6% baseline, and 17.2% withdraw it versus 2.7%. Among 1,656 high-uncertainty calls with return data, the median next-day return is -0.097 versus -0.072 baseline, and 37.0% beat, below the 39.5% baseline rate.
- High-uncertainty calls make up 11.3% of the 165,182-call corpus (18,704 calls), with a 95% interval of 11.2%-11.5%.
- Only 4.6% of high-uncertainty calls raise guidance versus 21.1% of baseline calls, while 32.6% lower guidance versus 11.6%.
- Stress language scores 3.66 on high-uncertainty calls versus 2.43 baseline, and confidence falls to 5.96 from 7.21.
- The 2020 trend peak shows 30.59% of calls scoring high on uncertainty, and the median next-day return for these calls is -0.097 versus -0.072 baseline.
1Introduction
Earnings calls are where companies are asked to describe an uncertain future in public, and most of the time management teams succeed in projecting calm. But a minority of calls visibly crack: hedging increases, answers get longer and less specific, and the promotion-heavy tone of a typical call gives way to something more defensive. Identifying what those calls look like, how common they are, and how their guidance behavior differs from the average call matters to anyone who reads transcripts for a living. This study examines 18,704 calls scoring 6 or higher on a 0-9 uncertainty meter, drawn from 165,182 calls between 1990 and 2026, and compares their language, guidance actions, and next-day returns to the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 uncertainty meter (n = 18,704; 11.3% of the reference set, 95% Wilson interval 11.2%–11.5%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile of high-uncertainty calls is distinctive: stress rises 1.23 points above baseline (3.66 vs 2.43), evasion 0.61 (3.31 vs 2.70), while confidence drops 1.26 (5.96 vs 7.21), promotion falls 0.90, and specificity slips 0.25. Guidance actions diverge most: 32.6% lower guidance (vs 11.6% baseline), 17.2% withdraw it (vs 2.7%), and only 4.6% raise it (vs 21.1%). The trend series peaks in 2020 at 30.59% of calls, far above 2018's 6.09% low. Returns are modestly weaker: median -0.097 vs -0.072 baseline, with 37.0% beating versus 39.5% baseline. The strongest overrepresented narrative is 'The Question Left Hanging' at 1.6x, while 'Guidance Worth Underwriting' appears at only 0.49x.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.10 | 6.86 | +0.24 |
| Evasion | 3.31 | 2.70 | +0.61 |
| Specificity | 7.31 | 7.56 | -0.25 |
| Stress | 3.66 | 2.43 | +1.23 |
| Promotion | 4.15 | 5.05 | -0.90 |
| Confidence | 5.96 | 7.21 | -1.26 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 4.6% | 21.1% |
| Maintained | 26.8% | 48.8% |
| Lowered | 32.6% | 11.6% |
| Withdrawn | 17.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Question Left Hanging | 1.60× | 76.9% | 48.0% |
| Results Worse Than Direction | 1.45× | 74.4% | 51.1% |
| Underused Fixed Costs | 1.45× | 60.5% | 41.6% |
| Scale-Dependent Advantage Claims | 1.37× | 15.1% | 11.1% |
| When the CFO Dominates | 1.32× | 18.6% | 14.1% |
| Guidance Worth Underwriting | 0.49× | 34.9% | 71.5% |
| Skeptic Reassured | 0.51× | 34.0% | 66.4% |
| Volume About to Step Up | 0.54× | 15.3% | 28.5% |
| Deferred Revenue Growing | 0.61× | 5.4% | 8.9% |
| Calls That Resolve Doubts | 0.64× | 50.5% | 79.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -9.7% | -7.2% |
| Interquartile range | -32.1% to +11.4% | — |
| Share beating SPY | 37.0% (95% CI 35%–39%) | 39.5% |
| Observations | 1,656 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| CNC | Q2 2025 | 2025-07-25 | F |
| MTH | Q2 2025 | 2025-07-25 | C |
| WF | Q2 2025 | 2025-07-25 | C |
| WZZAF | Q1 2026 | 2025-07-25 | F |
| MLLGF | Q2 2025 | 2025-07-25 | C+ |
| RNECF | Q2 2025 | 2025-07-25 | D |
| INTC | Q2 2025 | 2025-07-24 | D |
| SAM | Q2 2025 | 2025-07-24 | D |
4Discussion
A careful reader should conclude that calls scoring high on the uncertainty meter are rarer than most readers might guess, cluster around guidance cuts and withdrawals, and use measurably more stress and evasion language. They should not conclude that uncertainty scores cause weaker returns, that the 37.0% beat rate versus 39.5% baseline is exploitable, or that 2020's 30.59% peak predicts anything about future crises. The returns sample covers only 1,656 of these calls against a 22,449-call baseline, and differences of this size can arise from which companies face turbulent quarters in the first place. The patterns describe what these calls look like, not what they foretell.
5Limitations
The uncertainty score and behavioral profiles come from AI-read fields, which are noisy and can misclassify tone, hedging, or stress on any individual call. The returns comparison rests on 1,656 high-uncertainty calls within a 22,449-call base skewed toward liquid names, so coverage is not representative of all listed firms. Our own forward tests falsified directional prediction from these signals. Additionally, large language models partially remember famous stocks' histories, contaminating any backtest that reuses transcripts those models may have seen during training. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.