Research › Call Signals
Artul.ai Research LibraryStudy No. 13Call SignalsUpdated 2026-08-28

Everything Is Fine: Calls That Read Rehearsed

By Artul.ai Research Group · n = 63,830 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls that the model flagged as reading rehearsed, using a corpus of 165,182 calls from 1990 to 2026, of which 38.64% met the flag (95% CI 38.41%-38.88%). Rehearsed-flagged calls score lower on candor (6.52 vs 6.86) and specificity (7.38 vs 7.56), and higher on promotion (5.52 vs 5.05), evasion (2.81 vs 2.70), and stress (2.49 vs 2.43). The share has climbed from 32.37% of calls in 2020 to 45.01% in 2025. Overrepresented segments include Scale-Dependent Advantage Claims (1.7x lift) and When the CFO Dominates (1.34x), while Pricing Recovering appears at 0.68x. Among 7,895 calls with one-day returns, the median move was -0.089% versus -0.072% for the 22,449-call base, with 37.23% beating versus 39.47%.

Key findings
  • 38.64% of the 165,182 calls in the corpus were flagged as reading rehearsed (95% CI 38.41%-38.88%).
  • Rehearsed-flagged calls show a promotion score of 5.52 versus 5.05 for other calls, a +0.47 delta, alongside lower candor (6.52 vs 6.86) and specificity (7.38 vs 7.56).
  • The rehearsed share rose from 32.37% of calls in 2020 to 45.01% in 2025, with 18,899 calls flagged in 2022 and 17,639 in 2024.
  • Among 7,895 flagged calls with returns data, the median one-day move was -0.089% versus -0.072% for the base, and 37.23% beat versus a base rate of 39.47%.

1Introduction

Anyone who listens to earnings calls long enough develops an ear for the polished ones: the smooth transitions, the pre-packaged soundbites, the answers that arrive before the question finishes. Whether that ear tracks anything measurable is an open question. If rehearsed delivery is a meaningful marker, it should show up in how those calls score on candor, promotion, and stress dimensions, in which conversational patterns cluster around them, and in how their guidance behavior compares to the rest of the corpus. This study examines the 38,630 calls, out of 165,182, that the model answered YES to the battery item "Calls That Read Rehearsed," profiling their language scores, segment lifts, guidance mix, yearly prevalence from 2015 to 2025, and one-day returns.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Calls That Read Rehearsed" (n = 63,830; 38.6% of the reference set, 95% Wilson interval 38.4%–38.9%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Rehearsed-flagged calls look polished rather than forthcoming: promotion runs 5.52 versus 5.05 elsewhere (+0.47), while candor (6.52 vs 6.86) and specificity (7.38 vs 7.56) sit lower. Guidance is modestly less favorable: 18.53% raised guidance versus 21.05% for other calls, and 9.49% lowered it versus 11.56%. The strongest segment lift is Scale-Dependent Advantage Claims at 1.7x (18.78% vs 11.08%), followed by The Finished-Story Tell (1.37x) and When the CFO Dominates (1.34x); Pricing Recovering is underrepresented at 0.68x. The annual share climbed from 34.85% in 2017 to a 2020 dip of 32.37%, then rose steadily to 45.01% in 2025. One-day returns skew slightly negative: a median of -0.089% versus -0.072%, with beats at 37.23% versus 39.47%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.526.86-0.34
Evasion2.812.70+0.11
Specificity7.387.56-0.18
Stress2.492.43+0.06
Promotion5.525.05+0.47
Confidence7.267.21+0.04
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised18.5%21.1%
Maintained48.6%48.8%
Lowered9.5%11.6%
Withdrawn1.8%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims1.70×18.8%11.1%
The Finished-Story Tell1.37×6.0%4.4%
When the CFO Dominates1.34×18.9%14.1%
A Tiny Fraction of the Market1.32×39.5%30.0%
Pricing Recovering0.68×14.7%21.5%
201538.04%
201635.06%
201734.85%
201837.15%
201938.69%
202032.37%
202139.37%
202241.11%
202342.18%
202443.53%
202545.01%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-8.9%-7.2%
Interquartile range-28.3% to +10.7%
Share beating SPY37.2% (95% CI 36%–38%)39.5%
Observations7,89522,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
USCBQ2 20252025-07-25B+
HCAQ2 20252025-07-25C
AONQ2 20252025-07-25C
BFHQ2 20252025-07-25B
FFICQ2 20252025-07-25B+
OMFQ2 20252025-07-25A
GBCIQ2 20252025-07-25A
LARKQ2 20252025-07-25B

4Discussion

A careful reader should treat this as a descriptive portrait, not a warning sign. Rehearsed-flagged calls genuinely differ on measured dimensions: more promotional language, less candor, a rising prevalence, and a slightly weaker guidance mix and beat rate. But none of these differences establishes that rehearsal causes market outcomes, that the flag identifies deceptive management, or that the pattern is tradable. The returns gap is small and the samples differ in composition. What the data support is a consistent association between the flag and a particular communication style, one that has become more common over the past decade.

5Limitations

The flag comes from AI-read fields, which are noisy and reflect model interpretation rather than verified rehearsal. The returns sample covers only 7,895 flagged calls of a 22,449-call base skewed toward liquid names, so small-cap behavior is underrepresented. Our own forward tests falsified directional prediction on these signals, and LLMs partially remember famous stocks' histories, contaminating any backtest. Differences reported here are descriptive associations within the corpus and should not be read as causal effects or as a basis for trading decisions. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Everything Is Fine: Calls That Read Rehearsed.” Artul.ai Earnings-Call Research Library, Study No. 13. https://artul.ai/research/calls-that-read-rehearsed-earnings-calls

Related studies

Ammunition, Not Smoke: Inside the Calls Where CriticKnow What You Know: Calls Where Confidence Matched tThe Question Left Hanging, Fewer and Farther BetweenThe Guidance Was Fine All Along: What Model-EndorsedTalking the Skeptic Down: 66% of Calls Leave the DouThe Admission Is Not the Confession: Calls Flagged a
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.