Research › Management Behavior
Artul.ai Research LibraryStudy No. 12Management BehaviorUpdated 2026-08-28

39% of earnings calls trip the complexity meter - and they raise guidance less often

By Artul.ai Research Group · n = 64,523 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls scoring 6 or higher on a 0-9 language complexity meter. Of 165,182 calls from 1990-2026, 64,523 (39.06%, 95% CI 38.83%-39.30%) qualify. High-complexity calls show +0.34 higher evasion and +0.32 higher stress, but -0.14 lower confidence and -0.11 lower candor than typical calls. Only 16.22% of them raised guidance versus 21.05% of other calls. Their next-day returns are weaker: a median of -0.0894 versus -0.0716, with 36.98% beating the market versus 39.47%. Complexity peaked at 44.62% of calls in 2015 and fell to 32.77% by 2025.

Key findings
  • 64,523 of 165,182 calls (39.06%) scored 6 or higher on the 0-9 complexity meter.
  • High-complexity calls score +0.34 higher on evasion and +0.32 higher on stress, but -0.14 lower on confidence than other calls.
  • 16.22% of high-complexity calls raised guidance, versus 21.05% of all other calls.
  • Calls following high-complexity language had a median next-day return of -0.0894 versus -0.0716 for all other calls, with 36.98% beating the market versus 39.47%.

1Introduction

Anyone who reads earnings transcripts knows the feeling of a call that sounds elaborate but says little. Complexity in management language is easy to notice and hard to quantify, which is why a simple 0-9 complexity meter is useful: it turns that impression into a number that can be measured across decades of calls. This study asks what distinguishes the 39.06% of calls that score 6 or higher - how their candor, evasion, stress, and confidence differ, how often their companies raise guidance, and how their next-day returns compare with the rest of the corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 complexity meter (n = 64,523; 39.1% of the reference set, 95% Wilson interval 38.8%–39.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

High-complexity calls read differently: evasion is +0.34 higher and stress +0.32 higher, while confidence (-0.14), candor (-0.11), and specificity (-0.08) all sit below the corpus norm. The guidance pattern points the same direction - 16.22% raised guidance versus 21.05% elsewhere, and 12.70% lowered it versus 11.56%. Returns echo the tone gap: median next-day return -0.0894 versus -0.0716, and a 36.98% market-beating rate versus 39.47%. One phrase is overrepresented by 1.58x: 'Scale-Dependent Advantage Claims.' The trend is also striking - complexity fell from 44.62% of calls in 2015 to 32.77% in 2025, with a sharp drop around 2020 (35.91%).

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.756.86-0.11
Evasion3.042.70+0.34
Specificity7.477.56-0.08
Stress2.752.43+0.32
Promotion5.105.05+0.05
Confidence7.077.21-0.14
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised16.2%21.1%
Maintained49.7%48.8%
Lowered12.7%11.6%
Withdrawn2.3%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims1.58×17.5%11.1%
The Question Left Hanging1.31×62.8%48.0%
201544.62%
201643.90%
201743.71%
201843.02%
201942.69%
202035.91%
202134.24%
202236.82%
202337.31%
202435.98%
202532.77%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-8.9%-7.2%
Interquartile range-27.9% to +10.1%
Share beating SPY37.0% (95% CI 36%–38%)39.5%
Observations8,65922,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
DOCQ2 20252025-07-25C
HCAQ2 20252025-07-25C
CNCQ2 20252025-07-25F
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B
FLGQ2 20252025-07-25B
TNETQ2 20252025-07-25C+
DBOEYQ2 20252025-07-25B+

4Discussion

A careful reader should conclude that high-complexity calls co-occur with more evasive language, less raised guidance, and slightly weaker next-day returns. That is a pattern of association, not proof that complexity causes anything - complexity may simply travel with harder quarters, tougher questions, or stressed sectors. The return gap (-0.0894 vs -0.0716 median) is descriptive, not a signal to trade on. The right takeaway is that language complexity is a measurable lens on management posture, one worth tracking alongside fundamentals rather than instead of them.

5Limitations

Complexity, candor, and stress scores come from AI-read transcript fields and are noisy by nature. The returns comparison covers 8,659 calls within a base of 22,449, skewed toward liquid names, so medians and beat rates may not generalize. Our own forward tests falsified directional prediction from these features - the return gaps shown here are historical descriptions only. Finally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of language-based scoring against realized outcomes. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “39% of earnings calls trip the complexity meter - and they raise guidance less often.” Artul.ai Earnings-Call Research Library, Study No. 12. https://artul.ai/research/businesses-too-complex-to-explain

Related studies

97.97% of 161,829 earnings calls score high on speci96.2% of earnings calls score candid - yet median 1-95.5% of earnings calls since 1990 clear the 6-pointHalf of 165,182 earnings calls score low on evasion 37.3% of earnings calls score as promotional - and 2Only 11.3% of earnings calls score high on uncertain
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.