Specificity Is Next to Godliness: 161,829 Earnings Calls That Answer the Question
This study asks what distinguishes earnings calls that are highly specific—scoring 6 or higher on Artul.ai's 0–9 specificity meter. Of 165,182 calls from 1990 to 2026, 161,829 (97.97%) meet the bar. High-specificity calls show modestly higher candor (6.89 vs 6.86) and confidence (7.23 vs 7.21), and slightly lower evasion (2.67 vs 2.70). They raise guidance somewhat more often (21.40% vs 21.05%). Among 22,298 calls with measured post-call returns, the median return is -0.0712, essentially matching the full sample's -0.0716, and 39.51% beat the comparison threshold versus 39.47% overall. Specificity is pervasive, but its measured outcomes are strikingly ordinary.
- 97.97% of 165,182 calls (161,829) score 6 or higher on the 0–9 specificity meter.
- High-specificity calls show higher candor (6.89 vs 6.86) and specificity (7.63 vs 7.56) with lower evasion (2.67 vs 2.70).
- Guidance is raised on 21.40% of high-specificity calls versus 21.05% of all calls.
- Median post-call return is -0.0712 for high-specificity calls versus -0.0716 for the full sample, with 39.51% versus 39.47% beating.
1Introduction
Earnings-call watchers often treat specificity—concrete numbers, named drivers, direct answers—as the mark of a management team worth trusting. Vague language is assumed to hide something. Artul.ai's specificity meter, which scores calls from 0 to 9, offers a way to test how common that virtue really is and what accompanies it. This study examines 165,182 earnings calls from 1990 through 2026, profiling the 161,829 that score 6 or higher on the meter. We compare their candor, evasion, guidance behavior, and post-call returns against the full corpus, asking whether specificity coincides with noticeably different disclosure patterns or outcomes.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 specificity meter (n = 161,829; 98.0% of the reference set, 95% Wilson interval 97.9%–98.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Specificity is nearly universal: 97.97% of calls clear the 6-point bar, with the annual share ranging from 98.63% (2023) down to 91.17% in 2025, where the sample is partial. High-specificity calls edge the corpus on candor (6.89 vs 6.86), specificity (7.63 vs 7.56), and confidence (7.23 vs 7.21), while evasion is marginally lower (2.67 vs 2.70). Guidance is raised more often (21.40% vs 21.05%) and withdrawn slightly less (2.63% vs 2.66%). The returns picture is flat: median post-call return of -0.0712 versus -0.0716 overall, and a 39.51% versus 39.47% beat rate—differences too small to interpret.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.89 | 6.86 | +0.03 |
| Evasion | 2.67 | 2.70 | -0.03 |
| Specificity | 7.63 | 7.56 | +0.07 |
| Stress | 2.42 | 2.43 | -0.01 |
| Promotion | 5.05 | 5.05 | -0.00 |
| Confidence | 7.23 | 7.21 | +0.02 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 21.4% | 21.1% |
| Maintained | 49.3% | 48.8% |
| Lowered | 11.7% | 11.6% |
| Withdrawn | 2.6% | 2.7% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -7.1% | -7.2% |
| Interquartile range | -26.0% to +11.9% | — |
| Share beating SPY | 39.5% (95% CI 39%–40%) | 39.5% |
| Observations | 22,298 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| HCA | Q2 2025 | 2025-07-25 | C |
| AON | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
4Discussion
The honest takeaway is that scoring 6 or higher on specificity describes almost every call in the corpus, so it separates very little. The behavioral deltas—candor up 0.03, evasion down 0.03, raised guidance up about a third of a percentage point—are real in the data but tiny. Returns are effectively indistinguishable between the two groups. A careful reader should conclude that specificity, at this threshold, is a normal feature of modern disclosure, not a marker of anything exceptional. No statement here implies that specific calls cause better or worse outcomes, or that the meter predicts anything.
5Limitations
Scores are produced by AI-read fields and inherit their noise; a 6 versus a 5 may reflect model quirks as much as management style. The returns sample covers 22,449 calls and skews toward liquid names, so return comparisons may not generalize. Our own forward tests falsified directional prediction from these features. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest built on these scores. All figures describe the sample as measured and support no causal or predictive claims. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.