Only 1.75% of 165,182 earnings calls show high stress - and 32.5% of them lowered guidance
This study examines 2,893 earnings calls from a 165,182-call corpus spanning 1990-2026 that score 6 or higher on a 0-9 stress meter. High-stress calls are rare (1.75% of the corpus, 95% CI 1.69%-1.82%) and linguistically distinct: evasion scores run 4.25 versus 2.70 on typical calls, while confidence (5.64 vs 7.21) and specificity (6.84 vs 7.56) run lower. Guidance behavior diverges sharply: 32.46% of high-stress calls lowered guidance versus 11.56% of typical calls, and 10.54% withdrew guidance versus 2.66%. Among 232 high-stress calls with matched returns, the median next-day return was -0.118 versus -0.072 for the 22,449-call baseline. These are descriptive patterns, not predictive signals.
- High-stress calls make up 1.75% of the 165,182-call corpus (2,893 calls), with a 95% confidence interval of 1.69% to 1.82%.
- On high-stress calls, evasion scores average 4.25 versus 2.70 on typical calls, while confidence averages 5.64 versus 7.21.
- 32.46% of high-stress calls lowered guidance and 10.54% withdrew guidance, versus 11.56% and 2.66% on typical calls.
- The median next-day return after the 232 high-stress calls with matched returns was -0.118, versus -0.072 for the 22,449-call baseline.
1Introduction
Earnings calls are usually studied for what management says, but the tone under pressure may be just as distinctive. Calls where the stress meter reads 6 or higher on a 0-9 scale are uncommon - only 1.75% of 165,182 calls from 1990 through 2026 - which makes them a small but sharply defined slice of the disclosure record. If stress is visible in language, it should coincide with more evasion, less confidence, and more conservative guidance behavior. This study profiles those 2,893 calls across seven linguistic dimensions, their guidance actions, recurring behavioral markers, and next-day returns, describing how they differ from the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 stress meter (n = 2,893; 1.8% of the reference set, 95% Wilson interval 1.7%–1.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
High-stress calls score 3.78 points higher on stress (6.21 vs 2.43) and 1.55 points higher on evasion (4.25 vs 2.70), while running lower on confidence (5.64 vs 7.21), specificity (6.84 vs 7.56), and candor (6.46 vs 6.86). Guidance behavior diverges most: lowered guidance appears on 32.46% of high-stress calls versus 11.56% of typical calls, and withdrawn guidance on 10.54% versus 2.66%; raised guidance is rare (3.84% vs 21.05%). The most over-represented marker is 'Scale-Dependent Advantage Claims' at a 3.94 lift. The trend series shows annual stress rates falling from 3.07% in 2015 to 0.89% in 2021, then rebounding to 2.01% in 2022. Median next-day returns after high-stress calls were -0.118 versus -0.072 in the baseline.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.46 | 6.86 | -0.40 |
| Evasion | 4.25 | 2.70 | +1.55 |
| Specificity | 6.84 | 7.56 | -0.72 |
| Stress | 6.21 | 2.43 | +3.78 |
| Promotion | 4.83 | 5.05 | -0.22 |
| Confidence | 5.64 | 7.21 | -1.57 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 3.8% | 21.1% |
| Maintained | 24.4% | 48.8% |
| Lowered | 32.5% | 11.6% |
| Withdrawn | 10.5% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 3.94× | 43.7% | 11.1% |
| The Question Left Hanging | 2.05× | 98.5% | 48.0% |
| The Finished-Story Tell | 1.85× | 8.1% | 4.4% |
| Underused Fixed Costs | 1.83× | 76.2% | 41.6% |
| Results Worse Than Direction | 1.64× | 83.8% | 51.1% |
| Skeptic Reassured | 0.01× | 0.9% | 66.4% |
| Calls That Resolve Doubts | 0.10× | 7.8% | 79.5% |
| Guidance Worth Underwriting | 0.15× | 10.7% | 71.5% |
| Confidence Proportionate to Evidence | 0.38× | 33.6% | 87.6% |
| Deferred Revenue Growing | 0.62× | 5.5% | 8.9% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -11.8% | -7.2% |
| Interquartile range | -43.7% to +10.8% | — |
| Share beating SPY | 32.3% (95% CI 27%–39%) | 39.5% |
| Observations | 232 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOW | Q2 2025 | 2025-07-24 | D |
| CMA | Q2 2025 | 2025-07-18 | D |
| EDUC | Q4 2025 | 2025-07-07 | F |
| CJREF | Q3 2025 | 2025-06-26 | F |
| LVRO | Q2 2025 | 2025-06-20 | F |
| GRRR | Q1 2025 | 2025-06-18 | F |
| MCHOY | Q4 2025 | 2025-06-12 | F |
| NROM | Q1 2025 | 2025-06-10 | F |
4Discussion
A careful reader should conclude that calls scoring 6 or higher on the stress meter are rare and look different: more evasion, less confidence, and far more lowered or withdrawn guidance than typical calls. These are co-occurrences in the same calls, not evidence that stress causes weak outcomes or that spotting stress offers a trading edge. The returns gap (-0.118 vs -0.072 median) is a descriptive comparison, and our own forward tests falsified directional prediction. The reasonable takeaway is that stress-marked calls tend to accompany more cautious guidance behavior, nothing more.
5Limitations
The linguistic scores are produced by AI-read fields and are noisy; a 6-or-higher cutoff captures a fuzzy boundary, not a precise state. The returns sample covers 22,449 calls with matched prices, skewed toward liquid names, and only 232 high-stress calls have returns attached. Our own forward tests falsified directional prediction, so no timing or selection use is supported. LLMs partially remember famous stocks' histories, which can contaminate any backtest of these labels. The trend series also has uneven yearly sample sizes, from 3,483 calls in 2015 to 6,012 partial-year calls in 2025. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.