Calls that resolve doubts: 79.5% of 165,182 earnings calls, and 40.3% raised guidance
This study examines 165,182 earnings calls from 1990 to 2026 and isolates the 131,335 calls (79.5%, 95% CI 79.3%–79.7%) where the model answered YES to 'Calls That Resolve Doubts.' Compared with other calls, these show higher candor (6.99 vs 6.86), specificity (7.75 vs 7.56), and confidence (7.42 vs 7.21), with lower evasion (2.49 vs 2.70) and stress (2.12 vs 2.43). Guidance behavior differs too: 24.9% raised guidance versus 21.1% elsewhere. The over-represented topic 'The Question Left Hanging' appears at 0.72% versus 0.35% (lift 2.08). A median one-day return of -0.064% was slightly above the -0.072% baseline, and 40.3% beat, versus 39.5% baseline.
- 79.5% of 165,182 earnings calls (n=131,335) were classified as resolving doubts, with a 95% CI of 79.3% to 79.7%.
- Doubt-resolving calls show higher candor (6.99 vs 6.86), specificity (7.75 vs 7.56), and confidence (7.42 vs 7.21), and lower evasion (2.49 vs 2.70) and stress (2.12 vs 2.43).
- 24.9% of doubt-resolving calls raised guidance versus 21.1% of other calls, while 8.9% lowered guidance versus 11.6% elsewhere.
- In a 19,998-call returns subsample, the median one-day return was -0.064% versus -0.072% baseline, with 40.3% beating versus 39.5% baseline.
1Introduction
Earnings calls are the primary venue where management addresses analyst skepticism, and whether a call actually resolves doubts is a signal many listeners track informally. With 165,182 calls scored by a language model across 1990–2026, we can quantify how common this quality is and what linguistic and disclosure patterns accompany it. If doubt-resolving calls differ systematically in candor, evasion, and guidance behavior, that tells us how disclosure quality clusters. This study examines the 131,335 calls the model labeled YES to 'Calls That Resolve Doubts,' their language profile, guidance actions, topic over/under-representation, yearly trend, and one-day returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Calls That Resolve Doubts" (n = 131,335; 79.5% of the reference set, 95% Wilson interval 79.3%–79.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Doubt-resolving calls (79.5% of the corpus) skew toward clarity: candor 6.99 vs 6.86, specificity 7.75 vs 7.56, confidence 7.42 vs 7.21, with evasion at 2.49 vs 2.70 and stress at 2.12 vs 2.43. Guidance actions differ modestly: 24.9% raised guidance (vs 21.1%) and 52.3% maintained it (vs 48.8%), while only 8.9% lowered (vs 11.6%) and 2.2% withdrew (vs 2.7%). 'The Question Left Hanging' is over-represented at 0.72% vs 0.35% (lift 2.08), and 'Scale-Dependent Advantage Claims' at 0.33% vs 0.11% (lift 3.00). The share dipped to 74.25% in 2025 after peaking at 83.34% in 2020. Median one-day returns were -0.064% vs -0.072% baseline, with 40.3% beating vs 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.99 | 6.86 | +0.13 |
| Evasion | 2.49 | 2.70 | -0.21 |
| Specificity | 7.75 | 7.56 | +0.19 |
| Stress | 2.12 | 2.43 | -0.31 |
| Promotion | 5.01 | 5.05 | -0.04 |
| Confidence | 7.42 | 7.21 | +0.21 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 24.9% | 21.1% |
| Maintained | 52.3% | 48.8% |
| Lowered | 8.9% | 11.6% |
| Withdrawn | 2.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 0.33× | 3.6% | 11.1% |
| The Question Left Hanging | 0.72× | 34.7% | 48.0% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -6.4% | -7.2% |
| Interquartile range | -24.4% to +12.2% | — |
| Share beating SPY | 40.3% (95% CI 40%–41%) | 39.5% |
| Observations | 19,998 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| AON | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| OMF | Q2 2025 | 2025-07-25 | A |
4Discussion
A careful reader should conclude that doubt-resolving calls are the modal case and that they co-occur with somewhat more candor, specificity, confidence, and guidance raises. These are associations measured on the same calls, not evidence that resolving doubts causes better outcomes or that the language profile causes guidance actions. The returns differences are small (-0.064% vs -0.072% median; 40.3% vs 39.5% beat rate) and should not be read as an exploitable pattern. The 2025 figure (74.25%) rests on only 6,012 calls and may reflect partial-year data rather than a real decline.
5Limitations
The YES/NO labels come from AI-read fields and are noisy; a model's judgment of whether a call resolves doubts may not match a human analyst's. The returns subsample covers 22,449 calls and is skewed toward liquid names, so return statistics may not generalize. Our own forward tests falsified directional prediction, so nothing here should be treated as a trading edge. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of model-scored text. The 2025 trend point (74.25%, n=6,012) may reflect incomplete data. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.