Only 92 of 165,182 earnings calls scored 2 or lower on candor — what they share
This study profiles the rarest tail of Artul.ai's candor meter: earnings calls scoring 2 or lower on the 0–9 scale. Of 165,182 calls spanning 1990 to 2026, just 92 qualify — about 0.056% of the corpus (95% CI roughly 0.045% to 0.068%). These calls skew heavily toward 'The Question Left Hanging' (2.09x overrepresented), 'Calls That Read Rehearsed' (1.68x), and 'Founder-Led Companies' (1.46x). Themes like 'Calls That Resolve Doubts', 'Guidance Worth Underwriting', and 'Skeptic Reassured' appear in 0.0% of these calls. Guidance was raised on 2.17% of them. The share has risen sharply recently: 0.15% of calls in 2024 and 0.93% in 2025, versus about 0.01% in most prior years.
- Only 92 of 165,182 calls (about 0.056%) scored 2 or lower on the 0–9 candor meter.
- 'The Question Left Hanging' is 2.09x overrepresented in low-candor calls, with 'Calls That Read Rehearsed' at 1.68x and 'Founder-Led Companies' at 1.46x.
- Themes like 'Calls That Resolve Doubts', 'Guidance Worth Underwriting', and 'Skeptic Reassured' appear in 0.0% of these calls.
- The annual share of low-candor calls climbed from 0.15% in 2024 to 0.93% in 2025, after sitting near 0.01% from 2016 through 2023.
1Introduction
Earnings-call analysis usually focuses on what management says. The rarer signal may be what the call avoids. A call that scores 2 or lower on the 0–9 candor meter is, by construction, evasive, low on specificity, and short on confidence: average evasive behavior of 2.70 and specificity of 7.56, both far from typical calls. Because such calls are nearly absent from most of the 1990–2026 corpus — just 92 of 165,182 — any year where they cluster is worth a closer look. This study examines what these 92 calls have in common, how their guidance behavior compares, and how their frequency has shifted over time.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 candor meter (n = 92; 0.1% of the reference set, 95% Wilson interval 0.0%–0.1%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile separates these calls sharply from the baseline: promotion language averages 5.05 versus a delta of -3.83, and confidence 7.21 against a delta of -5.16, while candor itself sits at 0.12, a -6.74 delta. Theme overrepresentation points to rehearsed, question-deflecting calls — 'The Question Left Hanging' at 2.09x — often at founder-led companies (1.46x). Every 'underrepresented' theme, including 'The Finished-Story Tell' and 'The Hidden Segment', registers 0.0. On guidance, 2.17% raised and 1.09% maintained, with none lowered or withdrawn. The trend is the most striking result: a share near 0.01% for most of 2016–2023 jumps to 0.15% in 2024 and 0.93% in 2025 (6,012 calls scored).
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 0.12 | 6.86 | -6.74 |
| Evasion | 0.45 | 2.70 | -2.25 |
| Specificity | 0.34 | 7.56 | -7.22 |
| Stress | 0.66 | 2.43 | -1.77 |
| Promotion | 1.22 | 5.05 | -3.83 |
| Confidence | 2.05 | 7.21 | -5.16 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 2.2% | 21.1% |
| Maintained | 1.1% | 48.8% |
| Lowered | 0.0% | 11.6% |
| Withdrawn | 0.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| The Question Left Hanging | 2.09× | 100.0% | 48.0% |
| Calls That Read Rehearsed | 1.68× | 65.2% | 38.7% |
| Founder-Led Companies | 1.46× | 29.3% | 20.0% |
| Calls That Resolve Doubts | 0.00× | 0.0% | 79.5% |
| Guidance Worth Underwriting | 0.00× | 0.0% | 71.5% |
| Skeptic Reassured | 0.00× | 0.0% | 66.4% |
| The Finished-Story Tell | 0.00× | 0.0% | 4.4% |
| The Hidden Segment | 0.00× | 0.0% | 21.1% |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| CRDOF | Q1 2025 | 2025-05-29 | F |
| AMWD | Q4 2025 | 2025-05-29 | F |
| WRD | Q1 2025 | 2025-05-21 | F |
| XPEV | Q1 2025 | 2025-05-21 | F |
| DDD | Q1 2025 | 2025-05-13 | F |
| BLZE | Q1 2025 | 2025-05-11 | F |
| NOMD | Q1 2025 | 2025-05-10 | F |
| JYNT | Q1 2025 | 2025-05-10 | F |
4Discussion
A careful reader should conclude that calls scoring 2 or lower on the candor meter are extremely rare, disproportionately feature hanging questions and rehearsed delivery, and became noticeably more common in 2024–2025. That is a description of the corpus, not a diagnosis of any company. The 2025 figure rests on 6,012 scored calls, so the 0.93% share is based on a smaller base than earlier years. None of these statistics indicate how such calls relate to subsequent stock performance, and nothing here should be read as a signal to trade on. Treat the theme lifts as descriptive associations only.
5Limitations
Candor scores and theme tags are AI-read fields and inherit the noise of language models, including partial memorization of famous stocks' histories, which can contaminate any backtest. The returns sample covers 22,449 calls and is skewed toward liquid names, so results may not generalize to small caps. Most importantly, our own forward tests falsified directional prediction from these features; the 2024–2025 rise in low-candor calls is a frequency observation, not evidence of any market outcome. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.