Raise It and They Will Come: Anatomy of the Earnings-Call Guidance Raise
This study examines earnings calls where guidance was read as raised, drawing on 165,182 calls from 1990 to 2026. Raised-guidance calls account for 21.05% of the corpus (95% CI 20.86%-21.25%). These calls show a distinct language profile versus other calls: confidence 7.79 vs 7.21, specificity 7.82 vs 7.56, and stress 1.87 vs 2.43. The most overrepresented narrative is 'Deferred Revenue Growing' at 1.57x, while 'Scale-Dependent Advantage Claims' appears at only 0.42x. Among 6,688 raised-guidance calls with return data, the median forward return is -5.54%, compared with -7.16% for the base sample, and 41.81% beat the base median versus 39.47% overall.
- Raised-guidance calls make up 21.05% of the 165,182-call corpus, with a 95% CI of 20.86% to 21.25%.
- Language on raised-guidance calls is measurably different: confidence runs 7.79 vs 7.21 and stress 1.87 vs 2.43 on other calls.
- 'Deferred Revenue Growing' is the most overrepresented narrative at 1.57x, while 'Scale-Dependent Advantage Claims' is most underrepresented at 0.42x.
- Among the 6,688 raised-guidance calls with return data, the median forward return is -5.54% vs -7.16% for the base sample, with 41.81% beating the base median vs 39.47% overall.
1Introduction
For anyone who parses earnings calls, the guidance raise is the moment everyone waits for: management publicly upgrades its own forecast. But raises are only about one call in five in our corpus, and the language around them is worth understanding on its own terms. Do raised-guidance calls actually sound different from other calls? Which narratives cluster around them, and which are conspicuously absent? And what did forward returns look like afterward in this historical sample? This study describes 34,779 calls from 1990 to 2026 where guidance was read as raised, profiling their language, narrative composition, year-by-year frequency, and subsequent return distribution.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where guidance was read as raised (n = 34,779; 21.1% of the reference set, 95% Wilson interval 20.9%–21.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The language profile separates raised-guidance calls cleanly: confidence 7.79 vs 7.21, promotion 5.36 vs 5.05, and specificity 7.82 vs 7.56, with stress notably lower at 1.87 vs 2.43. Narratives tell a complementary story: concrete, verifiable framings like 'Deferred Revenue Growing' (1.57x) and 'Guidance Worth Underwriting' (1.27x) are overrepresented, while vague or defensive framings like 'Scale-Dependent Advantage Claims' (0.42x) and 'The Question Left Hanging' (0.61x) are rare. Frequency peaked at 31.45% of calls in 2021 before settling to 21.67% in 2024. Returns were less negative afterward (median -5.54% vs -7.16% base), with 41.81% beating the base median.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.90 | 6.86 | +0.04 |
| Evasion | 2.48 | 2.70 | -0.22 |
| Specificity | 7.82 | 7.56 | +0.27 |
| Stress | 1.87 | 2.43 | -0.56 |
| Promotion | 5.36 | 5.05 | +0.31 |
| Confidence | 7.79 | 7.21 | +0.57 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 100.0% | 21.1% |
| Maintained | 0.0% | 48.8% |
| Lowered | 0.0% | 11.6% |
| Withdrawn | 0.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Deferred Revenue Growing | 1.57× | 13.9% | 8.9% |
| Skeptic Reassured | 1.35× | 89.5% | 66.4% |
| Early Products Growing Fast | 1.29× | 49.7% | 38.5% |
| Guidance Worth Underwriting | 1.27× | 91.1% | 71.5% |
| Volume About to Step Up | 1.26× | 35.7% | 28.5% |
| Scale-Dependent Advantage Claims | 0.42× | 4.6% | 11.1% |
| Results Worse Than Direction | 0.51× | 26.1% | 51.1% |
| The Question Left Hanging | 0.61× | 29.3% | 48.0% |
| The Hidden Segment | 0.73× | 15.5% | 21.1% |
| Underused Fixed Costs | 0.74× | 30.8% | 41.6% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -5.5% | -7.2% |
| Interquartile range | -24.2% to +14.2% | — |
| Share beating SPY | 41.8% (95% CI 41%–43%) | 39.5% |
| Observations | 6,688 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| HCA | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| OMF | Q2 2025 | 2025-07-25 | A |
| GBCI | Q2 2025 | 2025-07-25 | A |
| MOG.A | Q3 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude that raised-guidance calls have a recognizable signature: higher confidence and specificity, lower stress, and a tilt toward concrete narratives over defensive ones. One should not conclude that the language caused better outcomes, that raises predict returns, or that the -5.54% median represents an opportunity. The return comparison is descriptive and historical, the samples differ in composition, and our own forward tests did not confirm any directional edge. Treat these figures as a map of how raise calls sound, not as a signal to trade on.
5Limitations
The language and narrative fields are AI-read and inherently noisy; classifications can be inconsistent across models and contexts. The returns sample covers 22,449 calls and skews toward liquid names, so the raised-guidance subset of 6,688 calls is not representative of the full universe. Our own forward tests falsified directional prediction from these features. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest-style comparison and inflate apparent patterns. All figures here should be read as descriptive history, not as generalizable or tradeable relationships. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.