Half of 165,182 earnings calls score low on evasion - and they raise guidance more
This study examines 86,188 earnings calls scoring 2 or lower on a 0-9 evasion meter, drawn from a corpus of 165,182 calls spanning 1990-2026, meaning 52.2% of all calls fall into this low-evasion group. Compared with the rest of the corpus, these calls show higher candor (7.12 vs 6.86), higher specificity (7.85 vs 7.56), lower stress (2.02 vs 2.43), and lower evasion (1.79 vs 2.70). They also raise guidance more often (24.1% vs 21.1%) and withdraw it less often (1.7% vs 2.7%). Among 10,966 calls with return data, the median follow-up return was -6.1% versus -7.2% for the 22,449-call baseline, with 40.6% beating versus 39.5%. These are descriptive differences, not causal or predictive claims.
- 52.2% of the 165,182-call corpus (86,188 calls) scores 2 or lower on the 0-9 evasion meter.
- Low-evasion calls show candor of 7.12 versus 6.86 and evasion of 1.79 versus 2.70 for the rest of the corpus.
- Guidance was raised on 24.1% of low-evasion calls versus 21.1% of others, and withdrawn on 1.7% versus 2.7%.
- Among 10,966 low-evasion calls with returns, the median return was -6.1% versus -7.2% for the 22,449-call baseline, with 40.6% beating versus 39.5%.
1Introduction
Evasion scoring gives analysts a way to quantify how directly management answers questions, but the behavior of highly forthcoming calls has rarely been profiled at scale. If roughly half of all earnings calls are candid by this measure, the group is large enough to shape how the whole corpus reads: its tone norms, its guidance habits, and its market context. For anyone who screens calls for signals of transparency, knowing what a low-evasion call actually looks like - and how it differs from everything else - is a useful baseline. This study profiles 86,188 calls scoring 2 or lower on the 0-9 evasion meter against the remaining corpus, covering language traits, guidance actions, flagged behaviors, and follow-up returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 evasion meter (n = 86,188; 52.2% of the reference set, 95% Wilson interval 51.9%–52.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Low-evasion calls differ from the rest of the corpus on every profiled trait: candor 7.12 vs 6.86, specificity 7.85 vs 7.56, confidence 7.36 vs 7.21, with lower stress (2.02 vs 2.43), lower evasion (1.79 vs 2.70), and slightly lower promotion (4.89 vs 5.05). Guidance actions skew constructive: raised on 24.1% vs 21.1% of other calls, withdrawn on 1.7% vs 2.7%. The two flagged behaviors are 'The Question Left Hanging' (55% vs 26.5%) and 'Scale-Dependent Advantage Claims' (60% vs 6.7%). The annual share of low-evasion calls rose from 48.0% in 2015 to 54.7% in 2024 and 58.0% in 2025 (partial year). Median follow-up returns were -6.1% vs -7.2% for the 22,449-call baseline, with 40.6% beating vs 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.12 | 6.86 | +0.26 |
| Evasion | 1.79 | 2.70 | -0.91 |
| Specificity | 7.85 | 7.56 | +0.29 |
| Stress | 2.02 | 2.43 | -0.41 |
| Promotion | 4.89 | 5.05 | -0.16 |
| Confidence | 7.36 | 7.21 | +0.15 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 24.1% | 21.1% |
| Maintained | 50.1% | 48.8% |
| Lowered | 10.4% | 11.6% |
| Withdrawn | 1.7% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Question Left Hanging | 0.55× | 26.5% | 48.0% |
| Scale-Dependent Advantage Claims | 0.60× | 6.7% | 11.1% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -6.1% | -7.2% |
| Interquartile range | -23.3% to +11.6% | — |
| Share beating SPY | 40.6% (95% CI 40%–42%) | 39.5% |
| Observations | 10,966 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| GBCI | Q2 2025 | 2025-07-25 | A |
| MOG.A | Q3 2025 | 2025-07-25 | B+ |
| FLG | Q2 2025 | 2025-07-25 | B |
| FRST | Q2 2025 | 2025-07-25 | A |
4Discussion
A careful reader should conclude that calls scoring 2 or lower on the evasion meter are, by construction and by the profiled traits, more direct: higher candor and specificity, lower stress, and modestly different guidance behavior. What should not be concluded is that candor causes better outcomes or that these calls are safer investments. The return gap is small in median terms, and the beat rates differ by about one point. The over-represented flags are interesting but descriptive. No claim here supports prediction, timing, or a trading edge; the study describes what low-evasion calls look like, not what they will do.
5Limitations
The evasion score and all profiled traits are AI-read fields, which are noisy and may encode model artifacts rather than pure human behavior. The returns comparison uses 22,449 calls, a subset skewed toward liquid names, so the -6.1% versus -7.2% medians may not generalize. Our own forward tests falsified directional prediction from these signals, so nothing here should be read as an edge. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of call-based signals. All differences reported are descriptive associations within this dataset and period only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.