When the model flags 'Volume About to Step Up' on 28.5% of calls, returns still trail
This study examines 47,036 earnings calls (out of 165,182, or 28.5%) where Artul.ai's model answered YES to "Volume About to Step Up," spanning 1990-2026. Flagged calls showed a notable behavioral profile: promotion language ran 5.58 vs 5.05 on unflagged calls (+0.53) and confidence 7.51 vs 7.21 (+0.30), while stress was lower (2.33 vs 2.43). Guidance outcomes were worse: 26.4% of flagged calls raised guidance vs 21.1% otherwise, and 8.7% lowered guidance vs 11.6%. Among 5,213 flagged calls with measured post-call moves, the median return was -10.4% vs -7.2% for the 22,449-call base, and 37.6% beat vs 39.5%. The data describe associations, not predictions.
- The model flagged 47,036 of 165,182 calls (28.5%, 95% CI 28.3%-28.7%) with YES on "Volume About to Step Up."
- Flagged calls scored higher on promotion (5.58 vs 5.05, +0.53) and confidence (7.51 vs 7.21, +0.30) than unflagged calls.
- Guidance was raised on 26.4% of flagged calls vs 21.1% of others, and lowered on 8.7% vs 11.6%.
- Median post-call return on flagged calls was -10.4% vs -7.2% in the 22,449-call base sample, with 37.6% beating vs 39.5%.
1Introduction
Volume expectations are a recurring theme on earnings calls: management teams often suggest trading activity, demand, or engagement is about to accelerate. With 47,036 of 165,182 calls (28.5%) drawing a YES from the model on "Volume About to Step Up," this is one of the more common forward-looking framings in the corpus. For analysts, the question is whether such language accompanies different behavior, guidance outcomes, and post-call results, or whether it is simply a stylistic staple. This study profiles flagged calls by year, language traits, guidance actions, and measured returns, without asserting any causal or predictive relationship.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Volume About to Step Up", tracked by year (n = 47,036; 28.5% of the reference set, 95% Wilson interval 28.3%–28.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
Flagged calls skew promotional: promotion scored 5.58 vs 5.05 (+0.53) and confidence 7.51 vs 7.21 (+0.30), while stress was slightly lower (2.33 vs 2.43). Over-represented phrases include "Scale-Dependent Advantage Claims" (1.61x) and "Underused Fixed Costs" (1.46x); "When the CFO Dominates" was under-represented at 0.65x. The annual flag rate climbed from 20.96% in 2015 to a peak of 37.74% in 2021 before settling near 25-29%. Outcomes did not follow the confident tone: 26.4% raised guidance vs 21.1% in the base, and the median return among 5,213 flagged calls was -10.4% vs -7.2% base, with 37.6% beating vs 39.5%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.78 | 6.86 | -0.08 |
| Evasion | 2.67 | 2.70 | -0.03 |
| Specificity | 7.58 | 7.56 | +0.02 |
| Stress | 2.33 | 2.43 | -0.10 |
| Promotion | 5.58 | 5.05 | +0.53 |
| Confidence | 7.51 | 7.21 | +0.30 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 26.4% | 21.1% |
| Maintained | 47.4% | 48.8% |
| Lowered | 8.7% | 11.6% |
| Withdrawn | 2.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 1.61× | 17.9% | 11.1% |
| Underused Fixed Costs | 1.46× | 60.6% | 41.6% |
| Early Products Growing Fast | 1.45× | 55.7% | 38.5% |
| A Tiny Fraction of the Market | 1.42× | 42.5% | 30.0% |
| Deferred Revenue Growing | 1.29× | 11.4% | 8.9% |
| When the CFO Dominates | 0.65× | 9.2% | 14.1% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -10.4% | -7.2% |
| Interquartile range | -32.0% to +12.2% | — |
| Share beating SPY | 37.6% (95% CI 36%–39%) | 39.5% |
| Observations | 5,213 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| FLG | Q2 2025 | 2025-07-25 | B |
| FRST | Q2 2025 | 2025-07-25 | A |
| HMDPF | Q2 2025 | 2025-07-25 | B |
| FFBC | Q2 2025 | 2025-07-25 | B+ |
| DBOEY | Q2 2025 | 2025-07-25 | B+ |
| TBBK | Q2 2025 | 2025-07-25 | C |
| SSB | Q2 2025 | 2025-07-25 | B+ |
| STEL | Q2 2025 | 2025-07-25 | B |
4Discussion
A careful reader should conclude that calls flagged for anticipated volume step-ups tend to carry more promotional and confident language, and that their measured guidance and return outcomes in this sample were somewhat weaker than the base. What one should not conclude is that the flag causes anything, that it predicts poor results, or that shorting flagged calls would have worked: these are descriptive associations in historical data, and our own forward tests have falsified directional prediction on similar signals. Treat the profile as context for reading calls, not as an edge.
5Limitations
AI-read fields are noisy and can mislabel tone or topics. The returns sample covers 5,213 flagged calls against a 22,449-call base skewed toward liquid names, so results may not generalize. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of model flags. All figures are descriptive of this corpus and support no causal or trading claims. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.