Backlog Is Growing, and So Is the Confidence: Earnings Call Signals, Measured
Across 165,182 earnings calls from 1990 to 2026, we identify the 57,249 calls (34.66%, 95% CI 34.43%-34.89%) where backlog was read as growing. Compared with the base population, these calls score higher on confidence (7.52 vs 7.21), specificity (7.72 vs 7.56), and promotion (5.24 vs 5.05), and lower on evasion (2.55 vs 2.70) and stress (2.16 vs 2.43). Guidance behavior skews positive: 29.49% raised guidance versus 21.06% in the base, while 8.66% lowered it versus 11.56%. Among the 8,171 calls with measurable forward returns, the median return was -4.96% versus -7.16% for the base, and 42.27% beat versus 39.47%. The phrase 'Deferred Revenue Growing' appears 2.16x more often than expected.
- 34.66% of calls (57,249 of 165,182) were read as backlog growing, with a 95% CI of 34.43% to 34.89%.
- These calls show higher confidence (7.52 vs 7.21) and lower stress (2.16 vs 2.43) than the base population.
- Guidance was raised on 29.49% of these calls versus 21.06% of base calls, and lowered on 8.66% versus 11.56%.
- Median forward returns were -4.96% versus -7.16% for the base, with 42.27% beating versus 39.47% (n = 8,171).
1Introduction
Few phrases on an earnings call carry as much weight as a management claim that backlog is growing. It sounds like a promise: orders are piling up, revenue is already booked, the future is arriving on schedule. Anyone who follows calls closely hears the phrase constantly, but its actual frequency and behavioral fingerprint are rarely quantified. Is a growing-backlog read accompanied by more confident delivery, more specific language, or better guidance behavior? Does it coincide with different subsequent returns? Using Artul.ai's AI-read fields across 165,182 calls from 1990 to 2026, this study examines the 57,249 calls where backlog was read as growing and compares them with the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where backlog was read as growing (n = 57,249; 34.7% of the reference set, 95% Wilson interval 34.4%–34.9%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Backlog-growth calls differ in tone before they differ in outcomes. Confidence is higher (7.52 vs 7.21) and stress lower (2.16 vs 2.43), while specificity (7.72 vs 7.56) and promotion (5.24 vs 5.05) both run above the base. Guidance behavior follows: 29.49% raised versus 21.06% in the base, and only 8.66% lowered versus 11.56%. The phrase 'Deferred Revenue Growing' appears 2.16x more often than its base rate would suggest, and 'Volume About to Step Up' 1.42x. Returns are less dramatic: a median of -4.96% versus -7.16%, and a beat rate of 42.27% versus 39.47% among the 8,171 calls with measurable outcomes. The trend series peaks in 2021 at 45.98% before easing to 33.48% in 2025.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.91 | 6.86 | +0.05 |
| Evasion | 2.55 | 2.70 | -0.15 |
| Specificity | 7.72 | 7.56 | +0.16 |
| Stress | 2.16 | 2.43 | -0.27 |
| Promotion | 5.24 | 5.05 | +0.19 |
| Confidence | 7.52 | 7.21 | +0.30 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 29.5% | 21.1% |
| Maintained | 49.6% | 48.8% |
| Lowered | 8.7% | 11.6% |
| Withdrawn | 1.8% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Deferred Revenue Growing | 2.16× | 19.1% | 8.9% |
| Volume About to Step Up | 1.42× | 40.4% | 28.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -5.0% | -7.2% |
| Interquartile range | -23.1% to +13.5% | — |
| Share beating SPY | 42.3% (95% CI 41%–43%) | 39.5% |
| Observations | 8,171 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| DOC | Q2 2025 | 2025-07-25 | C |
| USCB | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| MOG.A | Q3 2025 | 2025-07-25 | B+ |
| FLG | Q2 2025 | 2025-07-25 | B |
| FRST | Q2 2025 | 2025-07-25 | A |
| FFBC | Q2 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude that backlog-growth reads cluster with a distinctive communication style: more confident, more specific, less stressed, and more likely to raise guidance. That is an association measured on this corpus, not a mechanism. The returns gap (-4.96% vs -7.16% median; 42.27% vs 39.47% beat rate) is descriptive and drawn from a subset of 8,171 calls; it does not establish that backlog language causes better outcomes, nor that these differences would persist out of sample. Nothing here supports a trading rule. The most defensible takeaway is that the phrase marks a communication regime, not a guarantee.
5Limitations
The backlog field is produced by an LLM reading call transcripts, so misreads are possible and the 34.66% share carries that noise. The returns sample covers 22,449 calls and skews toward liquid names, so it is not representative of the full corpus. Artul.ai's own forward tests falsified directional prediction from these signals, and LLMs partially remember famous stocks' histories, contaminating any backtest. All comparisons here are descriptive associations within one dataset and should not be read as causal effects or as evidence of an exploitable edge. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.