87.8% of earnings calls contain critic ammunition - and guidance gets raised less often
We examine 145,051 earnings calls from 1990-2026 where Artul.ai's model answered YES to the battery item 'Critic Ammunition', drawn from a base of 165,182 calls (87.8% of the corpus, 95% CI 87.7%-88.0%). Calls in this group show slightly lower candor (6.83 vs 6.86), lower confidence (7.13 vs 7.21), and higher stress (2.58 vs 2.43) than the base. Guidance was raised on 18.6% of these calls versus 21.1% overall, while guidance was lowered on 13.0% versus 11.6%. Among 19,045 calls with return data, the median next-day return was -0.082 versus -0.072 for the base, and 38.4% beat versus 39.5%. The share of such calls peaked at 90.98% in 2015 and fell to 84.29% in 2021 before rising again to 89.99% in 2022.
- 87.8% of the 165,182-call corpus (145,051 calls) answered YES to 'Critic Ammunition'.
- Calls with critic ammunition raised guidance 18.6% of the time versus 21.1% for the full corpus, and lowered guidance 13.0% versus 11.6%.
- Profile scores differ modestly: stress 2.58 vs 2.43, confidence 7.13 vs 7.21, candor 6.83 vs 6.86.
- Median next-day return was -0.082% across 19,045 such calls versus -0.072% for 22,449 base calls, with 38.4% beating versus 39.5%.
1Introduction
Nearly nine in ten earnings calls in our library contain language the model flags as usable ammunition for a company critic - yet the market consequences of that flag are rarely examined. If critical content systematically accompanies weaker guidance posture, softer tone, or different post-call returns, the flag could serve as a useful descriptive lens on management communication. Conversely, if differences are small, the flag may be little more than background noise in nearly every transcript. This study describes 145,051 calls from 1990-2026 where the model answered YES to 'Critic Ammunition', comparing their language profiles, guidance actions, and post-call returns against the 165,182-call corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Critic Ammunition" (n = 145,051; 87.8% of the reference set, 95% Wilson interval 87.7%–88.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The flag is nearly universal: 87.8% of calls (145,051 of 165,182) contain critic ammunition. Language differences are small but consistent - stress is higher (2.58 vs 2.43), confidence lower (7.13 vs 7.21), and candor slightly lower (6.83 vs 6.86). Guidance posture differs more visibly: raised guidance appears on 18.6% of flagged calls versus 21.1% corpus-wide, while lowered guidance appears on 13.0% versus 11.6%. The annual share dips to 84.29% in 2021 and peaks at 90.98% in 2015, with 2022 at 89.99%. Post-call returns show a median of -0.082% versus -0.072% for the base, and 38.4% beat versus 39.5% - differences measured on 19,045 flagged calls with return data.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.83 | 6.86 | -0.03 |
| Evasion | 2.79 | 2.70 | +0.10 |
| Specificity | 7.51 | 7.56 | -0.04 |
| Stress | 2.58 | 2.43 | +0.15 |
| Promotion | 5.08 | 5.05 | +0.03 |
| Confidence | 7.13 | 7.21 | -0.09 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 18.6% | 21.1% |
| Maintained | 48.5% | 48.8% |
| Lowered | 13.0% | 11.6% |
| Withdrawn | 2.9% | 2.7% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -8.2% | -7.2% |
| Interquartile range | -27.5% to +11.2% | — |
| Share beating SPY | 38.4% (95% CI 38%–39%) | 39.5% |
| Observations | 19,045 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude only that calls containing critic ammunition tend, on average, to pair with slightly lower confidence, more stress, less frequent guidance raises, and marginally different returns. These are descriptive co-occurrences in a large sample, not evidence that the flagged language causes any outcome or that the flag predicts returns. The differences are small relative to the 87.8% prevalence, meaning the flag separates few calls from the corpus. No causal claim, forecast, or trading implication is supported by these statistics.
5Limitations
Language scores and battery answers are AI-read fields and inherently noisy; a YES on 'Critic Ammunition' reflects model judgment, not ground truth. The returns sample covers 22,449 calls, skewed toward liquid names, so return comparisons may not generalize. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of transcript-based measures. All figures here are descriptive only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.