Research › 20-Year Trends
Artul.ai Research LibraryStudy No. 3220-Year TrendsUpdated 2026-08-28

When the model flags 'Volume About to Step Up' on 28.5% of calls, returns still trail

By Artul.ai Research Group · n = 47,036 earnings calls · First published 2026-08-28
Abstract

This study examines 47,036 earnings calls (out of 165,182, or 28.5%) where Artul.ai's model answered YES to "Volume About to Step Up," spanning 1990-2026. Flagged calls showed a notable behavioral profile: promotion language ran 5.58 vs 5.05 on unflagged calls (+0.53) and confidence 7.51 vs 7.21 (+0.30), while stress was lower (2.33 vs 2.43). Guidance outcomes were worse: 26.4% of flagged calls raised guidance vs 21.1% otherwise, and 8.7% lowered guidance vs 11.6%. Among 5,213 flagged calls with measured post-call moves, the median return was -10.4% vs -7.2% for the 22,449-call base, and 37.6% beat vs 39.5%. The data describe associations, not predictions.

Key findings
  • The model flagged 47,036 of 165,182 calls (28.5%, 95% CI 28.3%-28.7%) with YES on "Volume About to Step Up."
  • Flagged calls scored higher on promotion (5.58 vs 5.05, +0.53) and confidence (7.51 vs 7.21, +0.30) than unflagged calls.
  • Guidance was raised on 26.4% of flagged calls vs 21.1% of others, and lowered on 8.7% vs 11.6%.
  • Median post-call return on flagged calls was -10.4% vs -7.2% in the 22,449-call base sample, with 37.6% beating vs 39.5%.

1Introduction

Volume expectations are a recurring theme on earnings calls: management teams often suggest trading activity, demand, or engagement is about to accelerate. With 47,036 of 165,182 calls (28.5%) drawing a YES from the model on "Volume About to Step Up," this is one of the more common forward-looking framings in the corpus. For analysts, the question is whether such language accompanies different behavior, guidance outcomes, and post-call results, or whether it is simply a stylistic staple. This study profiles flagged calls by year, language traits, guidance actions, and measured returns, without asserting any causal or predictive relationship.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Volume About to Step Up", tracked by year (n = 47,036; 28.5% of the reference set, 95% Wilson interval 28.3%–28.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Flagged calls skew promotional: promotion scored 5.58 vs 5.05 (+0.53) and confidence 7.51 vs 7.21 (+0.30), while stress was slightly lower (2.33 vs 2.43). Over-represented phrases include "Scale-Dependent Advantage Claims" (1.61x) and "Underused Fixed Costs" (1.46x); "When the CFO Dominates" was under-represented at 0.65x. The annual flag rate climbed from 20.96% in 2015 to a peak of 37.74% in 2021 before settling near 25-29%. Outcomes did not follow the confident tone: 26.4% raised guidance vs 21.1% in the base, and the median return among 5,213 flagged calls was -10.4% vs -7.2% base, with 37.6% beating vs 39.5%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.786.86-0.08
Evasion2.672.70-0.03
Specificity7.587.56+0.02
Stress2.332.43-0.10
Promotion5.585.05+0.53
Confidence7.517.21+0.30
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised26.4%21.1%
Maintained47.4%48.8%
Lowered8.7%11.6%
Withdrawn2.0%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims1.61×17.9%11.1%
Underused Fixed Costs1.46×60.6%41.6%
Early Products Growing Fast1.45×55.7%38.5%
A Tiny Fraction of the Market1.42×42.5%30.0%
Deferred Revenue Growing1.29×11.4%8.9%
When the CFO Dominates0.65×9.2%14.1%
201520.96%
201625.05%
201727.60%
201827.25%
201926.13%
202028.46%
202137.74%
202228.85%
202328.02%
202428.97%
202525.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-10.4%-7.2%
Interquartile range-32.0% to +12.2%
Share beating SPY37.6% (95% CI 36%–39%)39.5%
Observations5,21322,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
FLGQ2 20252025-07-25B
FRSTQ2 20252025-07-25A
HMDPFQ2 20252025-07-25B
FFBCQ2 20252025-07-25B+
DBOEYQ2 20252025-07-25B+
TBBKQ2 20252025-07-25C
SSBQ2 20252025-07-25B+
STELQ2 20252025-07-25B

4Discussion

A careful reader should conclude that calls flagged for anticipated volume step-ups tend to carry more promotional and confident language, and that their measured guidance and return outcomes in this sample were somewhat weaker than the base. What one should not conclude is that the flag causes anything, that it predicts poor results, or that shorting flagged calls would have worked: these are descriptive associations in historical data, and our own forward tests have falsified directional prediction on similar signals. Treat the profile as context for reading calls, not as an edge.

5Limitations

AI-read fields are noisy and can mislabel tone or topics. The returns sample covers 5,213 flagged calls against a 22,449-call base skewed toward liquid names, so results may not generalize. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of model flags. All figures are descriptive of this corpus and support no causal or trading claims. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “When the model flags 'Volume About to Step Up' on 28.5% of calls, returns still trail.” Artul.ai Earnings-Call Research Library, Study No. 32. https://artul.ai/research/how-often-is-volume-really-about-to-step-up

Related studies

87.8% of earnings calls contain critic ammunition - 66.4% of earnings calls left the model's skeptic reaWhen a Call Ends With a Question Left Hanging: 47.9538.6% of 165,182 earnings calls read rehearsed to ouFounder-led calls: 20.0% of the corpus, but only 41.When the CFO Dominates: 14.1% of 165,182 Calls, With
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.