The Story That Ends Too Neatly: A Study of 'The Finished-Story Tell' on Earnings Calls
We examined 7,247 earnings calls (4.39% of a 165,182-call corpus spanning 1990-2026) where Artul.ai's model answered YES to the battery item 'The Finished-Story Tell' — calls whose narrative wraps up with suspicious polish. These calls skew more specific (7.7 vs 7.56), more promotion-focused (5.13 vs 5.05), and slightly more stressed (2.53 vs 2.43), while reading as marginally less evasive (2.66 vs 2.7). Guidance-wise, 16.0% raised versus 21.1% of baseline calls, and 51.7% maintained guidance. Calls reading rehearsed over-index (1.38x), and CFO dominance lifts to 1.27x. Median forward returns on 919 flagged calls were -0.082 versus -0.072 baseline.
- Flagged calls make up 4.39% of the corpus (7,247 of 165,182 calls, 1990-2026).
- Calls That Read Rehearsed over-index at 1.38x and When the CFO Dominates at 1.27x among finished-story calls.
- Guidance raises appear on 16.0% of flagged calls versus 21.1% of the baseline; 51.7% maintained versus 48.8%.
- Median forward returns on 919 flagged calls were -0.082 versus -0.072 for 22,449 baseline calls, with 38.6% beating versus 39.5% baseline.
1Introduction
Every investor knows the feeling of a call that ends too cleanly: every question answered, every arc closed, no loose threads left for the analyst to pull. On earnings calls, an overly complete narrative may signal preparation, polish, or a story being managed rather than reported. Artul.ai's behavioral battery includes an item called 'The Finished-Story Tell' that flags exactly this pattern. We examined 7,247 such calls out of 165,182 total (4.39%) from 1990 to 2026, profiling their language, guidance behavior, contextual over- and under-indexing, and forward returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "The Finished-Story Tell" (n = 7,247; 4.4% of the reference set, 95% Wilson interval 4.3%–4.5%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Finished-story calls read slightly more specific (7.7 vs 7.56) and promotion-heavy (5.13 vs 5.05), with modestly higher stress (2.53 vs 2.43) and near-identical confidence. Context matters: they over-index when Calls That Read Rehearsed (1.38x) and When the CFO Dominates (1.27x), and under-index among Deferred Revenue Growing (0.68x), A Tiny Fraction of the Market (0.71x), and Founder-Led Companies (0.73x). Guidance raises are rarer (16.0% vs 21.1%) while maintained guidance is more common (51.7% vs 48.8%). Annual flag rates ranged from 3.78% (2020) to 6.89% (2015), settling at 4.29% in 2024. Median forward returns were -0.082 versus -0.072 baseline.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.90 | 6.86 | +0.04 |
| Evasion | 2.66 | 2.70 | -0.04 |
| Specificity | 7.70 | 7.56 | +0.14 |
| Stress | 2.53 | 2.43 | +0.10 |
| Promotion | 5.13 | 5.05 | +0.08 |
| Confidence | 7.24 | 7.21 | +0.02 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 16.0% | 21.1% |
| Maintained | 51.7% | 48.8% |
| Lowered | 12.3% | 11.6% |
| Withdrawn | 2.7% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Calls That Read Rehearsed | 1.38× | 53.5% | 38.7% |
| When the CFO Dominates | 1.27× | 17.9% | 14.1% |
| Deferred Revenue Growing | 0.68× | 6.0% | 8.9% |
| A Tiny Fraction of the Market | 0.71× | 21.3% | 30.0% |
| Founder-Led Companies | 0.73× | 14.6% | 20.0% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -8.2% | -7.2% |
| Interquartile range | -26.2% to +10.5% | — |
| Share beating SPY | 38.6% (95% CI 36%–42%) | 39.5% |
| Observations | 919 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| NWG | Q2 2025 | 2025-07-25 | B+ |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| CHTR | Q2 2025 | 2025-07-25 | C+ |
| SSB | Q2 2025 | 2025-07-25 | B+ |
| AUB | Q2 2025 | 2025-07-25 | B+ |
| WZZAF | Q1 2026 | 2025-07-25 | F |
| LGDDF | Q2 2025 | 2025-07-24 | C+ |
| BHLB | Q2 2025 | 2025-07-24 | B+ |
4Discussion
A finished-story flag co-occurs with rehearsed delivery and CFO-heavy calls, and with fewer guidance raises — but co-occurrence is not cause. The language profile differences (specificity 7.7 vs 7.56, stress 2.53 vs 2.43) are small in absolute terms. The returns gap (-0.082 vs -0.072 median) is modest and the beat rate is essentially flat (38.6% vs 39.5%). A careful reader should treat this as a descriptive portrait of a narrative style, not a signal that polished storytelling predicts anything about stock performance.
5Limitations
All fields are AI-read and noisy; flag rates and profile scores carry annotation error. The returns sample covers 22,449 calls skewed toward liquid names, so results may not generalize broadly. Our own forward tests falsified directional prediction, and no finding here should be read as an edge. Finally, LLMs partially remember famous stocks' history, contaminating any backtest — flagged-call returns may reflect memorized narratives rather than genuine behavioral signal. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.