On 14.1% of calls the CFO dominates — and raised guidance runs 17.7% vs 21.1%
What changes when the CFO dominates an earnings call? Across 165,182 calls from 1990 to 2026, the model answered YES for 23,336 calls (14.1%). These calls are measurably less promotional: promotion 4.46 vs 5.05 and confidence 6.95 vs 7.21, with higher specificity (7.71 vs 7.56) and stress (2.58 vs 2.43). Guidance leans conservative — raised on 17.7% of CFO-dominated calls vs 21.1% overall, maintained 53.0% vs 48.8%. Scripted-language markers over-index at 1.34x ('Calls That Read Rehearsed'), while hype staples like 'A Tiny Fraction of the Market' run at 0.64x. In the returns sample, the beat flag appears in 42.7% of 3,556 flagged calls vs 39.5% of 22,449 base calls, with median return -4.75% vs -7.16%.
- CFO-dominated calls make up 23,336 of 165,182 calls (14.1%) and score lower on promotion (4.46 vs 5.05) and confidence (6.95 vs 7.21) but higher on specificity (7.71 vs 7.56).
- Raised guidance appears on 17.7% of CFO-dominated calls vs 21.1% overall, while maintained guidance is more common (53.0% vs 48.8%) and lowered guidance slightly more common (13.1% vs 11.6%).
- Scripted-language markers over-index, with 'Calls That Read Rehearsed' at 1.34x the base rate (51.8% vs 38.7%), while hype phrases like 'A Tiny Fraction of the Market' under-index at 0.64x and 'Scale-Dependent Advantage Claims' at 0.62x.
- In the 3,556-call returns sample, the beat flag rate is 42.7% (95% CI 41.1% to 44.3%) vs 39.5% for 22,449 base calls, and the median return is -4.75% vs -7.16%.
1Introduction
Anyone who parses earnings calls knows the CEO-CFO dynamic matters: the chief financial officer owns the numbers, and who carries the call shapes what gets said. If CFO-dominated calls are less promotional and more scripted, listeners who lean on management tone may be reading a different signal than they assume, and teams scoring call language should know which features move with the format. The pattern is common — 14.1% of all calls. This study examines those 23,336 calls within a 165,182-call corpus spanning 1990-2026, where the model answered YES to the battery item 'When the CFO Dominates,' comparing their language profile, guidance mix, phrase rates, annual prevalence, and returns sample against the full corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "When the CFO Dominates" (n = 23,336; 14.1% of the reference set, 95% Wilson interval 14.0%–14.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile stands out: promotion runs 4.46 vs 5.05 (-0.59) and confidence 6.95 vs 7.21 (-0.26), while specificity (7.71 vs 7.56, +0.15) and stress (2.58 vs 2.43, +0.15) both tick up — a numerate, tenser tone. The guidance mix differs too: raised 17.7% vs 21.1%, maintained 53.0% vs 48.8%, lowered 13.1% vs 11.6%. Phrase rates sharpen the picture: 'Calls That Read Rehearsed' over-indexes at 1.34x (51.8% vs 38.7%), while market-size hype runs cold — 'Scale-Dependent Advantage Claims' at 0.62x, 'A Tiny Fraction of the Market' at 0.64x, 'Volume About to Step Up' at 0.65x. Prevalence drifts from 16.42% in 2015 to 12.74% in 2025, bottoming at 11.72% in 2021. In returns, the beat flag runs 42.7% vs 39.5% and median return -4.75% vs -7.16% (n=3,556 vs 22,449).
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.94 | 6.86 | +0.08 |
| Evasion | 2.68 | 2.70 | -0.01 |
| Specificity | 7.71 | 7.56 | +0.15 |
| Stress | 2.58 | 2.43 | +0.15 |
| Promotion | 4.46 | 5.05 | -0.59 |
| Confidence | 6.95 | 7.21 | -0.26 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 17.7% | 21.1% |
| Maintained | 53.0% | 48.8% |
| Lowered | 13.1% | 11.6% |
| Withdrawn | 2.8% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Calls That Read Rehearsed | 1.34× | 51.8% | 38.7% |
| The Finished-Story Tell | 1.27× | 5.6% | 4.4% |
| Scale-Dependent Advantage Claims | 0.62× | 6.8% | 11.1% |
| A Tiny Fraction of the Market | 0.64× | 19.3% | 30.0% |
| Volume About to Step Up | 0.65× | 18.5% | 28.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -4.8% | -7.2% |
| Interquartile range | -23.6% to +14.0% | — |
| Share beating SPY | 42.7% (95% CI 41%–44%) | 39.5% |
| Observations | 3,556 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| GBCI | Q2 2025 | 2025-07-25 | A |
| UVE | Q2 2025 | 2025-07-25 | C+ |
| FLG | Q2 2025 | 2025-07-25 | B |
4Discussion
The safest reading is descriptive: calls where the CFO dominates sound different — less promotional, more specific, more rehearsed — and their guidance mix leans conservative. None of this shows the CFO causes the tone, or that tone causes anything. Companies may simply hand the CFO the microphone in quieter quarters, in regulated industries, or when the story is operational rather than visionary. The higher beat flag (42.7% vs 39.5%) and less negative median return (-4.75% vs -7.16%) are single-sample comparisons, not an edge to trade on. Read this as a catalog of how call character varies with who leads it.
5Limitations
The model's YES/NO reads are noisy, and any single battery item misclassifies calls. The returns comparison rests on 22,449 calls with available data, skewed toward liquid names, so the 42.7% vs 39.5% beat rates and return medians may not generalize. Our own forward tests falsified directional prediction from these features — nothing here is an edge. Finally, LLMs partially remember famous stocks' histories, contaminating any backtest that mixes model labels with known outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.