Know What You Know: Calls Where Confidence Matched the Evidence
This study examines earnings calls where a language model answered YES to the battery item 'Confidence Proportionate to Evidence'. Of 165,182 calls scored from 1990 to 2026, 144,714 (87.6%, 95% CI 87.4% to 87.8%) met that bar. Relative to the remainder, these calls show higher candor (6.97 vs 6.86), higher specificity (7.67 vs 7.56), and lower stress (2.28 vs 2.43). Guidance was raised on 22.8% of matched calls vs 21.1% of others. Among 21,399 calls with post-call returns, the median return was -0.067 vs -0.072 for the base set. One battery item, 'Scale-Dependent Advantage Claims', appeared below its expected rate.
- 144,714 of 165,182 calls (87.6%, CI 87.4%-87.8%) were flagged as having confidence proportionate to evidence.
- Matched calls scored higher on candor (6.97 vs 6.86) and specificity (7.67 vs 7.56), and lower on stress (2.28 vs 2.43).
- Guidance was raised on 22.8% of matched calls versus 21.1% of the rest; lowered guidance was slightly rarer (11.0% vs 11.6%).
- Among calls with post-call returns, the median return was -0.067 for matched calls versus -0.072 in the 22,449-call base set.
- 'Scale-Dependent Advantage Claims' appeared 38% below its expected rate on matched calls (observed 4.2% vs 11.1% expected).
1Introduction
Anyone who parses earnings calls knows the feeling: a management team radiating certainty that the transcript does not quite support. A model-scored battery offers a systematic way to flag the opposite case - calls where expressed confidence lines up with the evidence presented. Because 87.6% of the corpus passes this filter, the interesting question is less who passes than what distinguishes the passing calls from the rest. This study examines 144,714 such calls against the remaining 20,468, comparing language profiles, guidance actions, theme rates, and post-call returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Confidence Proportionate to Evidence" (n = 144,714; 87.6% of the reference set, 95% Wilson interval 87.4%–87.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Matched calls tilt toward plainer speech: candor runs 6.97 versus 6.86, specificity 7.67 versus 7.56, while promotion language (4.89 vs 5.05) and stress (2.28 vs 2.43) both run lower. Guidance actions lean modestly positive: 22.8% raised versus 21.1% elsewhere, and 50.8% maintained versus 48.8%. Only one theme deviated from its expected rate - 'Scale-Dependent Advantage Claims' at 38% below expectation (4.2% observed vs 11.1% expected). The annual share is fairly stable, from 87.14% in 2015 to 86.56% in 2025, with a 2020 peak of 90.5% and a 2022 trough of 85.86%. Median post-call returns were -0.067 for matched calls versus -0.072 for the base.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.97 | 6.86 | +0.11 |
| Evasion | 2.60 | 2.70 | -0.10 |
| Specificity | 7.67 | 7.56 | +0.11 |
| Stress | 2.28 | 2.43 | -0.15 |
| Promotion | 4.89 | 5.05 | -0.16 |
| Confidence | 7.25 | 7.21 | +0.04 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 22.8% | 21.1% |
| Maintained | 50.8% | 48.8% |
| Lowered | 11.0% | 11.6% |
| Withdrawn | 2.6% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 0.38× | 4.2% | 11.1% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -6.7% | -7.2% |
| Interquartile range | -25.1% to +12.1% | — |
| Share beating SPY | 40.0% (95% CI 39%–41%) | 39.5% |
| Observations | 21,399 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should treat these results as descriptive. Calls rated as proportionate in confidence also read as more candid and specific, less promotional, and slightly less stressful - a coherent profile, but a correlation, not proof that one causes the other. The single under-weighted theme suggests such calls rarely lean on scale-dependent advantage claims. Post-call returns are statistically indistinguishable in practical terms, and none of these numbers support trading on the flag. The 2020 peak and 2022 trough are worth noting but not interpreting as market signals.
5Limitations
Battery items are model-generated and noisy; the YES label reflects one reading of one prompt. The returns sample covers only 21,399 of the flagged calls (22,449 in the base) and skews toward liquid names, so price outcomes may not generalize. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. Corpus coverage also varies by year, with 2025 partial at 6,012 calls, so trend comparisons across periods are uneven. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.