Research › Business Verdicts
Artul.ai Research LibraryStudy No. 94Business VerdictsUpdated 2026-08-28

Steady As She Goes: What Happens When Guidance Holds The Line

By Artul.ai Research Group · n = 80,591 earnings calls · First published 2026-08-28
Abstract

This study examines 80,591 earnings calls out of a 165,182-call corpus spanning 1990-2026 where guidance was read as maintained - 48.79% of all calls (95% CI 48.55%-49.03%). Maintained-guidance calls show modestly better communication profiles than the base: confidence 7.30 vs 7.21, specificity 7.63 vs 7.56, and stress 2.31 vs 2.43, with evasion slightly lower at 2.66 vs 2.70. The only overrepresented language theme is 'Scale-Dependent Advantage Claims' (lift 0.73). Annual share of maintained guidance dropped sharply to 36.31% in 2020 before recovering to 52.54% by 2025. Among 11,343 calls with return data, the median follow-up return was -0.07%, with 38.73% beating the base rate's 39.47%.

Key findings
  • Maintained guidance appears in 48.79% of the 165,182-call corpus (80,591 calls), with a 95% CI of 48.55% to 49.03%.
  • Maintained-guidance calls score higher on confidence (7.30 vs 7.21) and specificity (7.63 vs 7.56) and lower on stress (2.31 vs 2.43) than the base corpus.
  • The share of calls with maintained guidance fell to 36.31% in 2020 (17,007 calls) before climbing back to 52.54% in 2025 (6,012 calls).
  • The only overrepresented language theme is 'Scale-Dependent Advantage Claims' with a lift of 0.73, and 38.73% of the 11,343 return-sample calls beat the base rate of 39.47%.
  • Median follow-up return for maintained-guidance calls is -0.074%, essentially matching the base corpus median of -0.072%.

1Introduction

Most guidance talk on earnings calls is not a raise or a cut - it is a careful act of holding position. Nearly half of the 165,182 calls in our corpus (48.79%, or 80,591 calls) were read as maintaining prior guidance, making this the single most common guidance posture management takes. Yet 'maintained' is underexamined compared with dramatic lowers and withdrawals, even though it carries its own language signature and its own historical pattern across market regimes, from the 2020 collapse in maintenance to the slow rebuild afterward. This study profiles those 80,591 calls: how their tone differs, which themes surface more often, how the practice has trended since 2015, and how returns compare.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where guidance was read as maintained (n = 80,591; 48.8% of the reference set, 95% Wilson interval 48.5%–49.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Maintained-guidance calls read slightly more assured than the corpus at large: confidence runs 7.30 versus 7.21, specificity 7.63 versus 7.56, and stress 2.31 versus 2.43, with evasion a touch lower at 2.66 versus 2.70. The sole overrepresented theme, 'Scale-Dependent Advantage Claims', carries a lift of 0.73. The annual trend is striking: maintenance share fell from 55.58% in 2015 to 36.31% in 2020, then recovered steadily - 42.74% in 2021, 46.89% in 2022, 50.98% in 2023, 51.94% in 2024, and 52.54% in 2025. Returns offer no clear separation: median follow-up return of -0.074% versus -0.072% for the base, with 38.73% beating the base's 39.47%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.886.86+0.02
Evasion2.662.70-0.03
Specificity7.637.56+0.07
Stress2.312.43-0.12
Promotion4.985.05-0.07
Confidence7.307.21+0.09
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised0.0%21.1%
Maintained100.0%48.8%
Lowered0.0%11.6%
Withdrawn0.0%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims0.73×8.1%11.1%
201555.58%
201652.10%
201751.56%
201851.72%
201952.61%
202036.31%
202142.74%
202246.89%
202350.98%
202451.94%
202552.54%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-7.4%-7.2%
Interquartile range-25.6% to +10.4%
Share beating SPY38.7% (95% CI 38%–40%)39.5%
Observations11,34322,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
DOCQ2 20252025-07-25C
USCBQ2 20252025-07-25B+
AONQ2 20252025-07-25C
AMSFQ2 20252025-07-25C+
VRTSQ2 20252025-07-25C+
HMDPFQ2 20252025-07-25B
ULHQ2 20252025-07-25C+
FFBCQ2 20252025-07-25B+

4Discussion

The right takeaway is descriptive, not predictive. Maintained-guidance calls sound marginally more confident and less stressed than average, but the deltas are small and measured by AI readers, not ground truth. The 2020 trough and subsequent recovery coincide with well-known macro events; association across years is not evidence that guidance posture causes anything. Likewise, the returns comparison is a wash by design of the data - a 38.73% beat rate against a 39.47% base is a difference, not an edge. Readers should treat this as a profile of a common communication pattern, not a signal to act on.

5Limitations

Candor, evasion, specificity, stress, promotion, and confidence are AI-read fields and inherently noisy, so small deltas may not survive measurement error. The returns sample covers 22,449 calls but is skewed toward liquid names, and our own forward tests falsified directional prediction on this data. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of language against returns. The annual trend also mixes changing corpus composition with changing practice. All figures describe the past; none license causal or forward-looking claims. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Steady As She Goes: What Happens When Guidance Holds The Line.” Artul.ai Earnings-Call Research Library, Study No. 94. https://artul.ai/research/when-guidance-just-holds-earnings-calls

Related studies

Margins Are Expanding, and So Is the Confidence: ReaCapex Up, Stress Down: Reading 64,183 Calls Where SpGuidance Is a Vibe: Calls Where Demand Reads as AcceBacklog Is Growing, and So Is the Confidence: EarninMargins Are Fine, Thanks for Asking: What ContractioShow Me the Money (Ask): Calls Where Pricing Read as
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.