The Guidance Goes Quiet: Language and Outcomes of Withdrawal Calls
This study examines 4,393 earnings calls where guidance was read as withdrawn — about 2.7% of a 165,182-call corpus spanning 1990 to 2026. On these calls, management language shifts in a consistent direction: stress runs 0.92 points above the corpus baseline (3.35 vs 2.43), evasion is 0.59 points higher (3.28 vs 2.70), and confidence drops 0.83 points (6.39 vs 7.21), while promotion falls 0.42 and specificity slips 0.25. Candor is modestly elevated at 7.09 vs 6.86. Post-call returns are available for 127 such calls: the median is -15.9%, versus -7.2% for the 22,449-call base sample, and 35.4% beat, versus 39.5% at baseline. The 2020 spike — 15.9% of that year's calls — shows how withdrawal clusters in crisis periods.
- Withdrawn-guidance calls make up 2.7% of the corpus (4,393 of 165,182 calls).
- Stress language is 0.92 points higher than baseline on these calls (3.35 vs 2.43), the largest profile gap.
- Confidence is 0.83 points lower (6.39 vs 7.21) and promotion is 0.42 points lower (4.63 vs 5.05) than the corpus baseline.
- The median post-call return on 127 withdrawal calls is -15.9%, versus -7.2% for the 22,449-call base sample.
- Withdrawal peaked at 15.9% of calls in 2020, far above every other year in the trend.
1Introduction
Few sentences on an earnings call land harder than an announcement that guidance is being withdrawn. It removes the one number Wall Street anchors on, and it typically arrives amid operational uncertainty rather than in spite of it. For analysts, the question is whether the surrounding call — the tone, the hedging, the language patterns — carries a recognizable signature, and whether the aftermath looks systematically different from a typical call. This study examines 4,393 calls where guidance was read as withdrawn, drawn from a corpus of 165,182 calls covering 1990 through 2026, comparing their language profiles, post-call returns, and thematic content against the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where guidance was read as withdrawn (n = 4,393; 2.7% of the reference set, 95% Wilson interval 2.6%–2.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The language profile of withdrawal calls tilts toward difficulty: stress is up 0.92 points (3.35 vs 2.43) and evasion up 0.59 (3.28 vs 2.70), while confidence is down 0.83 (6.39 vs 7.21) and promotion down 0.42 (4.63 vs 5.05). Notably, candor is slightly higher (7.09 vs 6.86) — these calls read as strained but not evasive-first. Themes overrepresented include 'The Question Left Hanging' (1.57x) and 'Results Worse Than Direction' (1.47x). Post-call outcomes skew weak: the median return is -15.9% versus -7.2% at baseline, with 35.4% beating versus 39.5% at baseline — a wide interquartile range (-39.2% to 7.6%) cautions against treating withdrawal as a uniform event.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.09 | 6.86 | +0.23 |
| Evasion | 3.28 | 2.70 | +0.59 |
| Specificity | 7.31 | 7.56 | -0.25 |
| Stress | 3.35 | 2.43 | +0.92 |
| Promotion | 4.63 | 5.05 | -0.42 |
| Confidence | 6.39 | 7.21 | -0.83 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 0.0% | 21.1% |
| Maintained | 0.0% | 48.8% |
| Lowered | 0.0% | 11.6% |
| Withdrawn | 100.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Underused Fixed Costs | 1.62× | 67.4% | 41.6% |
| The Question Left Hanging | 1.57× | 75.1% | 48.0% |
| Results Worse Than Direction | 1.47× | 75.0% | 51.1% |
| Guidance Worth Underwriting | 0.12× | 8.3% | 71.5% |
| Pricing Recovering | 0.58× | 12.4% | 21.5% |
| Calls That Read Rehearsed | 0.69× | 26.8% | 38.7% |
| Deferred Revenue Growing | 0.71× | 6.3% | 8.9% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -15.9% | -7.2% |
| Interquartile range | -39.2% to +7.6% | — |
| Share beating SPY | 35.4% (95% CI 28%–44%) | 39.5% |
| Observations | 127 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| WZZAF | Q1 2026 | 2025-07-25 | F |
| VICR | Q2 2025 | 2025-07-22 | F |
| AEHR | Q4 2025 | 2025-07-08 | C |
| AOUT | Q4 2025 | 2025-06-26 | C |
| FDX | Q4 2025 | 2025-06-24 | C |
| LVRO | Q2 2025 | 2025-06-20 | F |
| VNCE | Q1 2025 | 2025-06-17 | F |
| JILL | Q1 2025 | 2025-06-11 | F |
4Discussion
A careful reader should conclude that withdrawn-guidance calls carry a distinct observable profile: more stress and evasion, less confidence and promotion, weaker median post-call returns, and heavy clustering in 2020. What should not be concluded is that withdrawal itself causes poor outcomes, that the language gap implies intent or deception, or that any of these statistics predicts returns — the distributions overlap substantially, and the beat-rate difference (35.4% vs 39.5%) is modest. The elevated candor score suggests withdrawal often accompanies frank acknowledgment, not concealment. These are descriptive patterns measured after the fact; they describe what withdrawal calls look like, not what to trade on.
5Limitations
The language fields are AI-read and noisy, so profile deltas like the 0.92-point stress gap should be read as tendencies, not precise measurements. The returns sample is 22,449 calls skewed toward liquid names, and only 127 withdrawal calls have returns attached — a small, selection-prone subset. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. The 2020 spike reflects a unique macro episode rather than a repeatable pattern. Nothing here establishes causation, prediction, or a trading edge. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.