The Sound of One Hand Answering: Earnings Calls That Stay Vague
This study examines 431 earnings calls, out of a corpus of 165,182 calls spanning 1990 to 2026, that score 2 or lower on Artul.ai's 0-9 specificity meter — roughly 0.26% of all calls. These calls are strikingly quieter across the board: specificity runs 1.31 versus a 7.56 baseline (a gap of -6.25), candor 3.42 versus 6.86 (-3.44), and confidence 4.35 versus 7.21 (-2.86). Guidance behavior differs too: 21.05% of baseline calls raise guidance, while none of these low-specificity calls do, and only 0.93% maintain guidance versus 48.79% at baseline. The phenomenon is also recent: it appears in under 0.05% of calls through 2023 but reaches 4.87% of calls in 2025.
- Low-specificity calls show a specificity score of 1.31 against a baseline of 7.56, a gap of -6.25, the largest behavioral delta in the study.
- Candor (3.42 vs 6.86, -3.44), promotion (2.34 vs 5.05, -2.71), and confidence (4.35 vs 7.21, -2.86) are all sharply lower on these calls.
- None of the 431 low-specificity calls raised guidance, compared with 21.05% of baseline calls, and only 0.93% maintained guidance versus 48.79% at baseline.
- The share of low-specificity calls jumped from 0.01% in 2023 to 0.69% in 2024 and 4.87% in 2025.
1Introduction
Anyone who follows earnings calls knows the feeling: a long Q&A, polished executives, and somehow no concrete answer to anything. Most calls leave traces — numbers, commitments, direct answers. A small slice does not. Identifying what those calls look like, and how common they are over time, matters because specificity is one of the few things a listener can judge without a model: did management actually say anything testable? This study examines 431 calls scoring 2 or lower on Artul.ai's 0-9 specificity meter, drawn from a corpus of 165,182 calls covering 1990 through 2026, comparing their language profile, guidance behavior, and prevalence over time.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 specificity meter (n = 431; 0.3% of the reference set, 95% Wilson interval 0.2%–0.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The language deltas are the headline. Against a baseline of 7.56, these calls score 1.31 on specificity — a gap of -6.25 — and they lag on candor (3.42 vs 6.86), promotion (2.34 vs 5.05), stress (1.32 vs 2.43), and confidence (4.35 vs 7.21). Guidance behavior is similarly muted: 0% raise guidance versus 21.05% at baseline, and 0.93% maintain it versus 48.79%. Thematic over-representation is mild at best — 'Calls That Read Rehearsed' sits at 2.26% versus a 0.88% base — while themes like 'Calls That Resolve Doubts', 'Guidance Worth Underwriting', and 'Skeptic Reassured' never appear in this group. The trend is the surprise: near-zero shares through 2023, then 0.69% in 2024 and 4.87% in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 3.42 | 6.86 | -3.44 |
| Evasion | 0.30 | 2.70 | -2.40 |
| Specificity | 1.31 | 7.56 | -6.25 |
| Stress | 1.32 | 2.43 | -1.11 |
| Promotion | 2.34 | 5.05 | -2.71 |
| Confidence | 4.35 | 7.21 | -2.86 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 0.0% | 21.1% |
| Maintained | 0.9% | 48.8% |
| Lowered | 0.2% | 11.6% |
| Withdrawn | 0.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Calls That Read Rehearsed | 2.26× | 87.7% | 38.7% |
| The Question Left Hanging | 2.09× | 100.0% | 48.0% |
| Scale-Dependent Advantage Claims | 1.59× | 17.6% | 11.1% |
| Calls That Resolve Doubts | 0.00× | 0.0% | 79.5% |
| Guidance Worth Underwriting | 0.00× | 0.0% | 71.5% |
| Skeptic Reassured | 0.00× | 0.0% | 66.4% |
| Pricing Recovering | 0.02× | 0.5% | 21.5% |
| Deferred Revenue Growing | 0.03× | 0.2% | 8.9% |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| CRDOF | Q1 2025 | 2025-05-29 | F |
| AMWD | Q4 2025 | 2025-05-29 | F |
| REX | Q1 2025 | 2025-05-28 | F |
| TPTA | Q1 2025 | 2025-05-22 | F |
| WB | Q1 2025 | 2025-05-21 | F |
| WRD | Q1 2025 | 2025-05-21 | F |
| MDWD | Q1 2025 | 2025-05-21 | D |
| BIOX | Q3 2025 | 2025-05-21 | D |
4Discussion
A careful reader should conclude that calls scoring very low on specificity are rare, measurably quieter in candor, confidence, and promotion, and much more common in 2024-2025 than in any prior year. They should not conclude that low specificity causes anything — these are descriptive associations, not causal claims — or that such calls predict returns, guidance misses, or trouble ahead. The 2025 figure partly reflects a smaller recent sample of 6,012 calls. What the data shows is a distinctive, increasingly frequent communication style; what it means for outcomes is a separate, unanswered question.
5Limitations
The specificity score and all behavioral fields are AI-read and noisy; a 0-2 rating may misclassify calls that a human would score differently. Any returns comparison would rest on a sample of 22,449 calls skewed toward liquid names. Our own forward tests falsified directional prediction, so nothing here should be read as an edge. LLMs partially remember famous stocks' histories, which can contaminate any backtest of scoring rules. Finally, the 2025 sample covers only 6,012 calls, so the apparent surge in low-specificity calls carries wide uncertainty. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.