The Admission Is Not the Confession: Calls Flagged as 'Results Worse Than Direction'
This study examines 84,487 earnings calls (out of 165,182, or 51.15%) where the model answered YES to the battery item 'Results Worse Than Direction', across a 1990–2026 corpus. Flagged calls show lower candor (6.86 vs 6.86), higher stress (2.94 vs 2.43) and evasion (2.85 vs 2.70), and lower confidence (6.85 vs 7.21). Guidance behavior diverges sharply: 17.64% lowered guidance versus 11.56% in the base, and 10.74% raised it versus 21.05%. In the returns subsample (n = 9,167), the median return was -0.098 versus -0.072 in the base, with 37.08% beating versus 39.47%. Flagged calls over-index on 'Scale-Dependent Advantage Claims' (lift 1.70) and under-index on 'Skeptic Reassured' (0.69).
- 51.15% of 165,182 calls (84,487) were flagged YES on 'Results Worse Than Direction', with a 95% share interval of 50.91% to 51.39%.
- Flagged calls show higher stress (2.94 vs 2.43) and evasion (2.85 vs 2.70), but nearly identical candor (6.86 vs 6.86).
- Guidance diverges: 17.64% of flagged calls lowered guidance vs 11.56% of the base, while only 10.74% raised it vs 21.05%.
- In the returns subsample, flagged calls had a median return of -0.098 vs -0.072 for the base and a beat rate of 37.08% (CI 36.10%–38.07%) vs 39.47%.
1Introduction
Earnings calls rarely announce bad news outright; the damage usually shows up in tone, hedging, and what management chooses to lower rather than raise. A model flag of 'Results Worse Than Direction' is one attempt to capture that moment systematically: the call where the numbers or the narrative read as worse than the trajectory the company had set. For anyone who reads calls closely, the question is whether such a flag tracks measurable differences in language, guidance behavior, and market reaction, or whether it mostly restates the obvious. This study compares 84,487 flagged calls against the 165,182-call corpus to describe those differences.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Results Worse Than Direction" (n = 84,487; 51.1% of the reference set, 95% Wilson interval 50.9%–51.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Flagged calls are not more candid than the base (6.86 vs 6.86) but read as more strained: stress runs 2.94 vs 2.43, evasion 2.85 vs 2.70, and confidence drops from 7.21 to 6.85. Guidance tells a sharper story: 17.64% of flagged calls lowered guidance versus 11.56% of the base, and only 10.74% raised it versus 21.05%. Rhetorically, flagged calls over-index on 'Scale-Dependent Advantage Claims' (lift 1.70), 'The Hidden Segment' (1.56), and 'Underused Fixed Costs' (1.43), while 'Skeptic Reassured' appears at 0.69 lift. In the returns subsample, the median flagged-call return was -0.098 versus -0.072, with a beat rate of 37.08% against 39.47%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.86 | 6.86 | -0.00 |
| Evasion | 2.85 | 2.70 | +0.15 |
| Specificity | 7.42 | 7.56 | -0.14 |
| Stress | 2.94 | 2.43 | +0.52 |
| Promotion | 5.04 | 5.05 | -0.01 |
| Confidence | 6.85 | 7.21 | -0.37 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 10.7% | 21.1% |
| Maintained | 47.2% | 48.8% |
| Lowered | 17.6% | 11.6% |
| Withdrawn | 3.9% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 1.70× | 18.9% | 11.1% |
| The Hidden Segment | 1.56× | 32.9% | 21.1% |
| Underused Fixed Costs | 1.43× | 59.8% | 41.6% |
| The Question Left Hanging | 1.29× | 61.7% | 48.0% |
| Skeptic Reassured | 0.69× | 45.9% | 66.4% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -9.8% | -7.2% |
| Interquartile range | -30.1% to +10.5% | — |
| Share beating SPY | 37.1% (95% CI 36%–38%) | 39.5% |
| Observations | 9,167 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| GBCI | Q2 2025 | 2025-07-25 | A |
| UVE | Q2 2025 | 2025-07-25 | C+ |
| VRTS | Q2 2025 | 2025-07-25 | C+ |
| FLG | Q2 2025 | 2025-07-25 | B |
| FRST | Q2 2025 | 2025-07-25 | A |
| HMDPF | Q2 2025 | 2025-07-25 | B |
4Discussion
A careful reader should treat these as descriptive co-movements, not causes or signals. Flagged calls do coincide with more guidance cuts, more stressed language, and somewhat weaker subsequent returns, but none of these gaps establishes that the flag predicts anything or that tone drove the outcomes. The returns difference (-0.098 vs -0.072 median) is modest and measured on a skewed subsample. The safest conclusion is that the flag marks a recognizable cluster of language and guidance behavior, and nothing more directional than that.
5Limitations
The underlying fields are AI-read and inherently noisy, so tone deltas of a few tenths should not be over-read. The returns analysis covers only 22,449 calls with outcomes, skewed toward liquid names, and our own forward tests falsified directional prediction from these features. LLMs also partially remember famous stocks' histories, which can contaminate any backtest by leaking hindsight into scores. All figures here are descriptive of the corpus and carry no causal or predictive claim. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.