When backlog reads as shrinking: 7,699 calls, and guidance was lowered 28.7% of the time
This study examines 7,699 earnings calls out of 165,182 (4.7% of the 1990-2026 corpus) in which backlog was read as shrinking. These calls show lower confidence (6.47 vs 7.21) and less promotion (4.39 vs 5.05) than the base, with notably higher stress (3.07 vs 2.43). Guidance behavior diverges sharply: guidance was lowered on 28.7% of these calls versus 11.6% overall, and raised on only 10.6% versus 21.1%. Forward returns for the 1,046 calls with outcome data show a median of -9.9% versus -7.2% in the base. Language patterns differ too: themes like 'Underused Fixed Costs' (1.44x) appear more often, while 'Deferred Revenue Growing' (0.50x) appears half as often. These are descriptive associations, not causal claims.
- Guidance was lowered on 28.7% of calls where backlog was read as shrinking, versus 11.6% in the base corpus.
- Median forward returns on these calls were -9.9%, versus -7.2% for the base sample of 22,449 calls.
- Confidence scored 6.47 on these calls versus 7.21 baseline, and stress was higher at 3.07 versus 2.43.
- The theme 'Deferred Revenue Growing' appeared at 0.50x its baseline rate, the most underrepresented pattern in these calls.
1Introduction
Backlog is one of the most-watched disclosures on earnings calls, and how it is described can reshape how listeners hear everything else in the quarter. Calls where backlog is read as shrinking represent about 4.7% of the Artul.ai corpus of 165,182 calls from 1990 to 2026, making them a small but distinct slice of disclosure behavior. Understanding how tone, guidance actions, and subsequent return distributions differ on these calls is useful context for anyone who parses management language. This study describes the characteristics of these 7,699 calls: their tone profile, guidance actions, recurring themes, frequency over time, and the distribution of forward returns.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where backlog was read as shrinking (n = 7,699; 4.7% of the reference set, 95% Wilson interval 4.6%–4.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The tone profile shows these calls carry less promotion (4.39 vs 5.05) and lower confidence (6.47 vs 7.21), with stress elevated at 3.07 versus 2.43, while specificity is essentially unchanged at 7.56 versus 7.56. Guidance actions diverge most: lowered guidance on 28.7% of calls versus 11.6% baseline, raised guidance on 10.6% versus 21.1%, and withdrawn guidance on 6.8% versus 2.7%. Themes tell a consistent story: 'Underused Fixed Costs' (1.44x) and 'Results Worse Than Direction' (1.33x) are overrepresented, while 'Deferred Revenue Growing' (0.50x) and 'Volume About to Step Up' (0.58x) are underrepresented. The annual share ranged from 1.9% in 2021 to 7.26% in 2015.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 7.18 | 6.86 | +0.32 |
| Evasion | 2.74 | 2.70 | +0.04 |
| Specificity | 7.56 | 7.56 | +0.00 |
| Stress | 3.07 | 2.43 | +0.64 |
| Promotion | 4.39 | 5.05 | -0.66 |
| Confidence | 6.47 | 7.21 | -0.74 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 10.6% | 21.1% |
| Maintained | 37.9% | 48.8% |
| Lowered | 28.7% | 11.6% |
| Withdrawn | 6.8% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Underused Fixed Costs | 1.44× | 60.1% | 41.6% |
| Results Worse Than Direction | 1.33× | 68.0% | 51.1% |
| The Hidden Segment | 1.25× | 26.5% | 21.1% |
| Deferred Revenue Growing | 0.50× | 4.4% | 8.9% |
| Volume About to Step Up | 0.58× | 16.6% | 28.5% |
| Early Products Growing Fast | 0.67× | 25.9% | 38.5% |
| A Tiny Fraction of the Market | 0.68× | 20.4% | 30.0% |
| Pricing Recovering | 0.70× | 15.0% | 21.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -9.9% | -7.2% |
| Interquartile range | -30.5% to +10.9% | — |
| Share beating SPY | 37.0% (95% CI 34%–40%) | 39.5% |
| Observations | 1,046 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| HMDPF | Q2 2025 | 2025-07-25 | B |
| MTH | Q2 2025 | 2025-07-25 | C |
| WZZAF | Q1 2026 | 2025-07-25 | F |
| RGP | Q4 2025 | 2025-07-24 | F |
| DAIO | Q2 2025 | 2025-07-24 | D |
| FPH | Q2 2025 | 2025-07-24 | D |
| HZO | Q3 2025 | 2025-07-24 | D |
| JAKK | Q2 2025 | 2025-07-24 | D |
4Discussion
A careful reader should conclude that calls where backlog was read as shrinking are associated with a different disclosure posture: more lowered guidance, less promotional tone, and a forward return distribution whose median (-9.9%) sits below the base (-7.2%). What should not be concluded is that shrinking backlog language causes poor outcomes, predicts returns, or offers a trading edge. The return distributions overlap substantially (interquartile range -30.5% to 10.9%), and 37.0% of these calls still preceded a beat. These are descriptive patterns in historical data, not forecasts.
5Limitations
Backlog readings are AI-generated from call language and are noisy; misclassification is possible in both directions. The returns sample covers 1,046 calls against a base of 22,449, skewed toward liquid names, so results may not generalize. Artul.ai's own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, which can contaminate backtests of language-based measures. All figures describe the observed sample only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.