Margins Are Fine, Thanks for Asking: What Contraction Talk Sounds Like
We analyzed 55,774 earnings calls—33.8% of the 165,182-call corpus (1990-2026)—where margins were read as contracting. These calls carry a distinct vocal and verbal signature: confidence runs 6.71 versus 7.21 on the baseline, specificity 7.45 versus 7.56, and stress 3.02 versus 2.43. Guidance behavior diverges sharply: 20.8% of such calls lowered guidance versus 11.6% overall, and 4.6% withdrew it versus 2.7%. Returns after these calls are worse: median -8.95% against -7.16% for the base, with only 38.0% beating the baseline rate of 39.5%. The most overrepresented narrative pattern, 'Results Worse Than Direction,' appears 1.53x more often than expected.
- Margin-contraction calls make up 33.8% of the 165,182-call corpus (n=55,774).
- Guidance was lowered on 20.8% of these calls versus 11.6% overall, and withdrawn on 4.6% versus 2.7%.
- Confidence scores drop to 6.71 versus a 7.21 baseline, while stress rises to 3.02 versus 2.43.
- Median post-call returns are -8.95% versus -7.16% for the base sample of 22,449 calls.
1Introduction
Few phrases on an earnings call travel further than 'margins.' When a call signals margin contraction, it reshapes how every other answer in the transcript is heard—deflections sound evasive, optimism sounds forced, and guidance language gets parsed for damage. For analysts, these calls are high-stakes moments where tone and substance both shift. Yet the pattern itself is rarely quantified: how common is it, how does management behavior change, and what does the guidance record look like? This study examines 55,774 calls from a 165,182-call corpus spanning 1990-2026 where margins were read as contracting, profiling their language, guidance actions, and follow-on outcomes.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where margins was read as contracting (n = 55,774; 33.8% of the reference set, 95% Wilson interval 33.5%–34.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile is consistent with strain: confidence falls 0.51 points (6.71 vs 7.21), promotion framing drops 0.24, and stress rises 0.59 (3.02 vs 2.43), while evasion edges up 0.16. Guidance skews negative—20.8% lowered versus 11.6% baseline, 4.6% withdrawn versus 2.7%, and only 10.0% raised versus 21.1%. Narrative patterns amplify the theme: 'Results Worse Than Direction' runs at 1.53x expected frequency and 'Underused Fixed Costs' at 1.30x, while 'Skeptic Reassured' (0.69x) and 'Deferred Revenue Growing' (0.73x) are rare. Returns are modestly worse: median -8.95% versus -7.16%, with 38.0% beats versus 39.5%. The trend line fell from 40.0% in 2022 to 25.28% in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.95 | 6.86 | +0.09 |
| Evasion | 2.85 | 2.70 | +0.16 |
| Specificity | 7.45 | 7.56 | -0.11 |
| Stress | 3.02 | 2.43 | +0.59 |
| Promotion | 4.81 | 5.05 | -0.24 |
| Confidence | 6.71 | 7.21 | -0.51 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 10.0% | 21.1% |
| Maintained | 45.1% | 48.8% |
| Lowered | 20.8% | 11.6% |
| Withdrawn | 4.6% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Results Worse Than Direction | 1.53× | 78.4% | 51.1% |
| Scale-Dependent Advantage Claims | 1.37× | 15.1% | 11.1% |
| Underused Fixed Costs | 1.30× | 54.0% | 41.6% |
| The Hidden Segment | 1.25× | 26.5% | 21.1% |
| The Question Left Hanging | 1.25× | 60.2% | 48.0% |
| Skeptic Reassured | 0.69× | 45.9% | 66.4% |
| Deferred Revenue Growing | 0.73× | 6.5% | 8.9% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -9.0% | -7.2% |
| Interquartile range | -29.5% to +11.0% | — |
| Share beating SPY | 38.0% (95% CI 37%–39%) | 39.5% |
| Observations | 6,778 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| CNC | Q2 2025 | 2025-07-25 | F |
| UVE | Q2 2025 | 2025-07-25 | C+ |
| HMDPF | Q2 2025 | 2025-07-25 | B |
| ULH | Q2 2025 | 2025-07-25 | C+ |
| TNET | Q2 2025 | 2025-07-25 | C+ |
| MTH | Q2 2025 | 2025-07-25 | C |
| PUBL | Q2 2025 | 2025-07-25 | C |
| VWAGY | Q2 2025 | 2025-07-25 | C |
4Discussion
A careful reader should treat these findings as description, not diagnosis. Calls read as margin-contracting genuinely differ in tone, guidance behavior, and narrative texture—the 20.8%-versus-11.6% lowering gap and the 1.53x lift for 'Results Worse Than Direction' are large. But nothing here establishes that the language causes outcomes or predicts them; the return gap (-8.95% vs -7.16% median) is a measured association in a skewed sample. The declining trend since 2022 may reflect macro conditions as much as disclosure behavior. Read this as a map of how contraction talk sounds, not a signal.
5Limitations
The margin-contraction label is an AI-read field and inherits LLM noise, so some calls are surely misclassified. The returns comparison covers 6,778 of these calls against a base of 22,449—a sample skewed toward liquid, widely covered names, so outcomes may not generalize. Our own forward tests falsified directional prediction from these features, and LLMs partially remember famous stocks' histories, contaminating any backtest with hindsight leakage. Confidence intervals on shares are narrow only because the corpus is large; they say nothing about labeling accuracy. Treat all cross-group differences as descriptive. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.