Only 861 of 165,182 earnings calls score under 3 on complexity - why they stand out
This study examines a rare slice of earnings-call language: calls scoring 2 or lower on Artul.ai's 0-9 complexity meter. Among 165,182 calls from 1990-2026, only 861 (0.52%, 95% CI 0.49%-0.56%) qualify. Compared with the corpus baseline, these calls show lower evasion (2.19 vs 2.70), lower stress (2.15 vs 2.43), and lower promotion language (4.45 vs 5.05), with candor slightly higher (7.09 vs 6.86). Guidance is maintained far more often (37.3% vs 48.8% baseline implies less maintaining; raised guidance is nearly identical at 21.1% vs 21.1%). A forward-returns sample of 99 calls shows a median next-period return of -0.85%, versus -7.16% for the 22,449-call baseline, and 48.5% raised guidance versus 39.5% baseline. These are descriptive contrasts, not causal claims.
- Only 861 of 165,182 calls (0.52%) score 2 or lower on the 0-9 complexity meter.
- Simple-language calls show lower evasion (2.19 vs 2.70) and lower stress (2.15 vs 2.43) than the corpus baseline.
- Raised guidance occurred in 48.5% of the 99-call returns sample versus 39.5% for the 22,449-call baseline.
- The median forward return in the simple-language sample was -0.85% versus -7.16% for the baseline, with wide dispersion (Q1 -20.4%, Q3 24.1%).
1Introduction
Earnings calls are famously dense: hedged statements, promotional framing, and jargon can obscure what management actually knows. A small set of calls does the opposite, using plain, direct language. Identifying what characterizes these calls - and whether their tone aligns with guidance behavior and subsequent returns - is useful context for anyone who reads transcripts closely. This study examines every call in the Artul.ai corpus scoring 2 or lower on the 0-9 complexity meter, comparing their language profile, guidance actions, topic emphasis, and forward returns against the full 165,182-call baseline from 1990-2026.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 complexity meter (n = 861; 0.5% of the reference set, 95% Wilson interval 0.5%–0.6%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Simple-language calls are notably less evasive (2.19 vs 2.70) and less stressed (2.15 vs 2.43), with less promotion (4.45 vs 5.05) and marginally higher candor (7.09 vs 6.86). Their topic mix over-indexes on 'Volume About to Step Up' (46% vs 13.2%, lift 0.28) and 'Early Products Growing Fast' (50% vs 19.2%, lift 0.39), while 'The Finished-Story Tell' is rare (6.1% vs 2.7%). Guidance is maintained less often (37.3% vs 48.8%) and withdrawn more often (4.2% vs 2.7%). In the 99-call returns sample, the median forward return is -0.85% versus -7.16% baseline, though the mean (5.60%) sits well above the median, indicating skew. The simple-language share of calls peaked at 0.77 in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.09 | 6.86 | +0.23 |
| Evasion | 2.19 | 2.70 | -0.51 |
| Specificity | 7.63 | 7.56 | +0.07 |
| Stress | 2.15 | 2.43 | -0.28 |
| Promotion | 4.45 | 5.05 | -0.60 |
| Confidence | 7.17 | 7.21 | -0.04 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 21.1% | 21.1% |
| Maintained | 37.3% | 48.8% |
| Lowered | 9.6% | 11.6% |
| Withdrawn | 4.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Hidden Segment | 0.26× | 5.5% | 21.1% |
| Scale-Dependent Advantage Claims | 0.27× | 3.0% | 11.1% |
| Volume About to Step Up | 0.46× | 13.2% | 28.5% |
| Early Products Growing Fast | 0.50× | 19.2% | 38.5% |
| The Finished-Story Tell | 0.61× | 2.7% | 4.4% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -0.9% | -7.2% |
| Interquartile range | -20.4% to +24.1% | — |
| Share beating SPY | 48.5% (95% CI 39%–58%) | 39.5% |
| Observations | 99 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| VRSN | Q2 2025 | 2025-07-24 | A |
| WTBA | Q2 2025 | 2025-07-24 | B+ |
| ORLY | Q2 2025 | 2025-07-24 | B+ |
| ACU | Q2 2025 | 2025-07-23 | A |
| CPAC | Q2 2025 | 2025-07-23 | B |
| SMPL | Q3 2025 | 2025-07-10 | C |
| SWBI | Q4 2025 | 2025-06-18 | D |
| CANADA | Q1 2026 | 2025-06-18 | D |
4Discussion
A careful reader should conclude that calls in unusually plain language are measurably less evasive, less promotional, and more often accompanied by maintained-or-changed guidance rather than the baseline pattern. They should not conclude that simplicity causes better outcomes, that these calls predict returns, or that the median-return gap represents an exploitable difference. The returns sample is small (99 calls), the distribution is highly skewed (mean 5.60% vs median -0.85%), and the confidence interval on the beat rate (38.9%-58.2%) overlaps substantially with the baseline. Treat every contrast here as descriptive.
5Limitations
Language metrics are AI-read fields and carry label noise. The returns comparison uses 22,449 baseline calls skewed toward liquid names, and the simple-language sample is only 99 calls, so sampling error is large. Artul.ai's own forward tests falsified directional prediction from these features, so no edge should be assumed. Additionally, LLMs partially remember the documented history of famous stocks, which can contaminate any backtest of language scores against realized returns. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.