95.5% of earnings calls since 1990 clear the 6-point confidence bar - why so few don't
This study asks how common it is for earnings calls to score 6 or higher on Artul.ai's 0-9 confidence meter, and how those high-confidence calls differ from the rest. Across 165,182 calls from 1990 to 2026, 157,716 (95.5%, 95% CI 95.4%-95.6%) met the threshold. High-confidence calls skew slightly more candid (6.87 vs 6.86), specific (7.61 vs 7.56), and promotional (5.12 vs 5.05), with lower stress (2.34 vs 2.43). They raise guidance more often (22.0% vs 21.1%) and lower it less often (10.2% vs 11.6%). Among 22,449 calls with return data, the median return is -0.070 versus -0.072 for the base sample.
- 95.5% of 165,182 calls (157,716) score 6 or higher on the 0-9 confidence meter (95% CI: 95.4% to 95.6%).
- High-confidence calls show lower stress (2.34 vs 2.43) and higher specificity (7.61 vs 7.56) than the full corpus.
- Guidance is raised on 22.0% of high-confidence calls versus 21.1% overall, and lowered on 10.2% versus 11.6%.
- The share of calls clearing the bar peaked at 98.08% in 2021 and fell to 90.05% in 2025 (6,012 calls).
1Introduction
Anyone who reads earnings calls regularly knows most transcripts read as confident, but confidence meters are rarely checked against the full distribution of calls. If nearly every call scores above a given bar, the bar tells you little; if the share moves over time, it may track how management teams communicate. This study measures how often calls score 6 or higher on the 0-9 confidence meter across 165,182 calls spanning 1990 to 2026, how those calls differ on behavioral profiles like candor, evasion, and stress, and how their guidance actions and post-call returns compare with the broader sample.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 confidence meter (n = 157,716; 95.5% of the reference set, 95% Wilson interval 95.4%–95.6%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The 6+ bar captures 95.5% of calls, so the interesting variation sits in the remaining slice and in year-to-year movement. The share climbed from 92.59% in 2015 (3,483 calls) to a peak of 98.08% in 2021 (18,253 calls), then slid to 90.05% in 2025. High-confidence calls are marginally more candid (6.87 vs 6.86), specific (7.61 vs 7.56), and promotional (5.12 vs 5.05), and less stressed (2.34 vs 2.43). They raise guidance slightly more often (22.0% vs 21.1%) and lower it less often (10.2% vs 11.6%). Median post-call returns are nearly identical: -0.070 versus -0.072 for the base sample, with 39.6% beating versus 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.87 | 6.86 | +0.01 |
| Evasion | 2.67 | 2.70 | -0.03 |
| Specificity | 7.61 | 7.56 | +0.05 |
| Stress | 2.34 | 2.43 | -0.09 |
| Promotion | 5.12 | 5.05 | +0.07 |
| Confidence | 7.33 | 7.21 | +0.12 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 22.0% | 21.1% |
| Maintained | 50.2% | 48.8% |
| Lowered | 10.2% | 11.6% |
| Withdrawn | 2.3% | 2.7% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -7.0% | -7.2% |
| Interquartile range | -25.8% to +11.9% | — |
| Share beating SPY | 39.6% (95% CI 39%–40%) | 39.5% |
| Observations | 21,927 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
4Discussion
The main takeaway is that a 6-or-higher confidence score is the norm, not a signal: 95.5% of calls clear it, and the behavioral and guidance differences versus the full corpus are small fractions of a point. A careful reader should not conclude that high confidence causes better outcomes, that the 2025 dip predicts anything, or that the near-identical median returns imply the meter is meaningless. The year-over-year trend may reflect changes in disclosure norms, scoring models, or sample composition. These are descriptive associations measured on one scoring system at one threshold.
5Limitations
The behavioral fields are AI-read and noisy, so small deltas like 0.01 in candor should not be over-interpreted. The returns sample covers 22,449 calls, skewed toward liquid names, and our own forward tests falsified directional prediction from these scores. LLMs partially remember famous stocks' histories, which can contaminate any backtest of call-level scores. The 2025 figure also reflects a partial year of 6,012 calls, making trend comparisons fragile. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.