Research › Management Behavior
Artul.ai Research LibraryStudy No. 70Management BehaviorUpdated 2026-08-28

Only 1.75% of 165,182 earnings calls show high stress - and 32.5% of them lowered guidance

By Artul.ai Research Group · n = 2,893 earnings calls · First published 2026-08-28
Abstract

This study examines 2,893 earnings calls from a 165,182-call corpus spanning 1990-2026 that score 6 or higher on a 0-9 stress meter. High-stress calls are rare (1.75% of the corpus, 95% CI 1.69%-1.82%) and linguistically distinct: evasion scores run 4.25 versus 2.70 on typical calls, while confidence (5.64 vs 7.21) and specificity (6.84 vs 7.56) run lower. Guidance behavior diverges sharply: 32.46% of high-stress calls lowered guidance versus 11.56% of typical calls, and 10.54% withdrew guidance versus 2.66%. Among 232 high-stress calls with matched returns, the median next-day return was -0.118 versus -0.072 for the 22,449-call baseline. These are descriptive patterns, not predictive signals.

Key findings
  • High-stress calls make up 1.75% of the 165,182-call corpus (2,893 calls), with a 95% confidence interval of 1.69% to 1.82%.
  • On high-stress calls, evasion scores average 4.25 versus 2.70 on typical calls, while confidence averages 5.64 versus 7.21.
  • 32.46% of high-stress calls lowered guidance and 10.54% withdrew guidance, versus 11.56% and 2.66% on typical calls.
  • The median next-day return after the 232 high-stress calls with matched returns was -0.118, versus -0.072 for the 22,449-call baseline.

1Introduction

Earnings calls are usually studied for what management says, but the tone under pressure may be just as distinctive. Calls where the stress meter reads 6 or higher on a 0-9 scale are uncommon - only 1.75% of 165,182 calls from 1990 through 2026 - which makes them a small but sharply defined slice of the disclosure record. If stress is visible in language, it should coincide with more evasion, less confidence, and more conservative guidance behavior. This study profiles those 2,893 calls across seven linguistic dimensions, their guidance actions, recurring behavioral markers, and next-day returns, describing how they differ from the rest of the corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 stress meter (n = 2,893; 1.8% of the reference set, 95% Wilson interval 1.7%–1.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

High-stress calls score 3.78 points higher on stress (6.21 vs 2.43) and 1.55 points higher on evasion (4.25 vs 2.70), while running lower on confidence (5.64 vs 7.21), specificity (6.84 vs 7.56), and candor (6.46 vs 6.86). Guidance behavior diverges most: lowered guidance appears on 32.46% of high-stress calls versus 11.56% of typical calls, and withdrawn guidance on 10.54% versus 2.66%; raised guidance is rare (3.84% vs 21.05%). The most over-represented marker is 'Scale-Dependent Advantage Claims' at a 3.94 lift. The trend series shows annual stress rates falling from 3.07% in 2015 to 0.89% in 2021, then rebounding to 2.01% in 2022. Median next-day returns after high-stress calls were -0.118 versus -0.072 in the baseline.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.466.86-0.40
Evasion4.252.70+1.55
Specificity6.847.56-0.72
Stress6.212.43+3.78
Promotion4.835.05-0.22
Confidence5.647.21-1.57
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised3.8%21.1%
Maintained24.4%48.8%
Lowered32.5%11.6%
Withdrawn10.5%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims3.94×43.7%11.1%
The Question Left Hanging2.05×98.5%48.0%
The Finished-Story Tell1.85×8.1%4.4%
Underused Fixed Costs1.83×76.2%41.6%
Results Worse Than Direction1.64×83.8%51.1%
Skeptic Reassured0.01×0.9%66.4%
Calls That Resolve Doubts0.10×7.8%79.5%
Guidance Worth Underwriting0.15×10.7%71.5%
Confidence Proportionate to Evidence0.38×33.6%87.6%
Deferred Revenue Growing0.62×5.5%8.9%
20153.07%
20162.49%
20172.14%
20181.94%
20191.73%
20201.24%
20210.89%
20222.01%
20231.81%
20241.54%
20251.45%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-11.8%-7.2%
Interquartile range-43.7% to +10.8%
Share beating SPY32.3% (95% CI 27%–39%)39.5%
Observations23222,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
DOWQ2 20252025-07-24D
CMAQ2 20252025-07-18D
EDUCQ4 20252025-07-07F
CJREFQ3 20252025-06-26F
LVROQ2 20252025-06-20F
GRRRQ1 20252025-06-18F
MCHOYQ4 20252025-06-12F
NROMQ1 20252025-06-10F

4Discussion

A careful reader should conclude that calls scoring 6 or higher on the stress meter are rare and look different: more evasion, less confidence, and far more lowered or withdrawn guidance than typical calls. These are co-occurrences in the same calls, not evidence that stress causes weak outcomes or that spotting stress offers a trading edge. The returns gap (-0.118 vs -0.072 median) is a descriptive comparison, and our own forward tests falsified directional prediction. The reasonable takeaway is that stress-marked calls tend to accompany more cautious guidance behavior, nothing more.

5Limitations

The linguistic scores are produced by AI-read fields and are noisy; a 6-or-higher cutoff captures a fuzzy boundary, not a precise state. The returns sample covers 22,449 calls with matched prices, skewed toward liquid names, and only 232 high-stress calls have returns attached. Our own forward tests falsified directional prediction, so no timing or selection use is supported. LLMs partially remember famous stocks' histories, which can contaminate any backtest of these labels. The trend series also has uneven yearly sample sizes, from 3,483 calls in 2015 to 6,012 partial-year calls in 2025. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Only 1.75% of 165,182 earnings calls show high stress - and 32.5% of them lowered guidance.” Artul.ai Earnings-Call Research Library, Study No. 70. https://artul.ai/research/management-under-visible-stress

Related studies

97.97% of 161,829 earnings calls score high on speci96.2% of earnings calls score candid - yet median 1-95.5% of earnings calls since 1990 clear the 6-pointHalf of 165,182 earnings calls score low on evasion 39% of earnings calls trip the complexity meter - an37.3% of earnings calls score as promotional - and 2
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.