Everything Is Fine: Calls That Read Rehearsed
This study examines earnings calls that the model flagged as reading rehearsed, using a corpus of 165,182 calls from 1990 to 2026, of which 38.64% met the flag (95% CI 38.41%-38.88%). Rehearsed-flagged calls score lower on candor (6.52 vs 6.86) and specificity (7.38 vs 7.56), and higher on promotion (5.52 vs 5.05), evasion (2.81 vs 2.70), and stress (2.49 vs 2.43). The share has climbed from 32.37% of calls in 2020 to 45.01% in 2025. Overrepresented segments include Scale-Dependent Advantage Claims (1.7x lift) and When the CFO Dominates (1.34x), while Pricing Recovering appears at 0.68x. Among 7,895 calls with one-day returns, the median move was -0.089% versus -0.072% for the 22,449-call base, with 37.23% beating versus 39.47%.
- 38.64% of the 165,182 calls in the corpus were flagged as reading rehearsed (95% CI 38.41%-38.88%).
- Rehearsed-flagged calls show a promotion score of 5.52 versus 5.05 for other calls, a +0.47 delta, alongside lower candor (6.52 vs 6.86) and specificity (7.38 vs 7.56).
- The rehearsed share rose from 32.37% of calls in 2020 to 45.01% in 2025, with 18,899 calls flagged in 2022 and 17,639 in 2024.
- Among 7,895 flagged calls with returns data, the median one-day move was -0.089% versus -0.072% for the base, and 37.23% beat versus a base rate of 39.47%.
1Introduction
Anyone who listens to earnings calls long enough develops an ear for the polished ones: the smooth transitions, the pre-packaged soundbites, the answers that arrive before the question finishes. Whether that ear tracks anything measurable is an open question. If rehearsed delivery is a meaningful marker, it should show up in how those calls score on candor, promotion, and stress dimensions, in which conversational patterns cluster around them, and in how their guidance behavior compares to the rest of the corpus. This study examines the 38,630 calls, out of 165,182, that the model answered YES to the battery item "Calls That Read Rehearsed," profiling their language scores, segment lifts, guidance mix, yearly prevalence from 2015 to 2025, and one-day returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Calls That Read Rehearsed" (n = 63,830; 38.6% of the reference set, 95% Wilson interval 38.4%–38.9%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Rehearsed-flagged calls look polished rather than forthcoming: promotion runs 5.52 versus 5.05 elsewhere (+0.47), while candor (6.52 vs 6.86) and specificity (7.38 vs 7.56) sit lower. Guidance is modestly less favorable: 18.53% raised guidance versus 21.05% for other calls, and 9.49% lowered it versus 11.56%. The strongest segment lift is Scale-Dependent Advantage Claims at 1.7x (18.78% vs 11.08%), followed by The Finished-Story Tell (1.37x) and When the CFO Dominates (1.34x); Pricing Recovering is underrepresented at 0.68x. The annual share climbed from 34.85% in 2017 to a 2020 dip of 32.37%, then rose steadily to 45.01% in 2025. One-day returns skew slightly negative: a median of -0.089% versus -0.072%, with beats at 37.23% versus 39.47%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.52 | 6.86 | -0.34 |
| Evasion | 2.81 | 2.70 | +0.11 |
| Specificity | 7.38 | 7.56 | -0.18 |
| Stress | 2.49 | 2.43 | +0.06 |
| Promotion | 5.52 | 5.05 | +0.47 |
| Confidence | 7.26 | 7.21 | +0.04 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 18.5% | 21.1% |
| Maintained | 48.6% | 48.8% |
| Lowered | 9.5% | 11.6% |
| Withdrawn | 1.8% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 1.70× | 18.8% | 11.1% |
| The Finished-Story Tell | 1.37× | 6.0% | 4.4% |
| When the CFO Dominates | 1.34× | 18.9% | 14.1% |
| A Tiny Fraction of the Market | 1.32× | 39.5% | 30.0% |
| Pricing Recovering | 0.68× | 14.7% | 21.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -8.9% | -7.2% |
| Interquartile range | -28.3% to +10.7% | — |
| Share beating SPY | 37.2% (95% CI 36%–38%) | 39.5% |
| Observations | 7,895 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| OMF | Q2 2025 | 2025-07-25 | A |
| GBCI | Q2 2025 | 2025-07-25 | A |
| LARK | Q2 2025 | 2025-07-25 | B |
4Discussion
A careful reader should treat this as a descriptive portrait, not a warning sign. Rehearsed-flagged calls genuinely differ on measured dimensions: more promotional language, less candor, a rising prevalence, and a slightly weaker guidance mix and beat rate. But none of these differences establishes that rehearsal causes market outcomes, that the flag identifies deceptive management, or that the pattern is tradable. The returns gap is small and the samples differ in composition. What the data support is a consistent association between the flag and a particular communication style, one that has become more common over the past decade.
5Limitations
The flag comes from AI-read fields, which are noisy and reflect model interpretation rather than verified rehearsal. The returns sample covers only 7,895 flagged calls of a 22,449-call base skewed toward liquid names, so small-cap behavior is underrepresented. Our own forward tests falsified directional prediction on these signals, and LLMs partially remember famous stocks' histories, contaminating any backtest. Differences reported here are descriptive associations within the corpus and should not be read as causal effects or as a basis for trading decisions. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.