Research › Management Behavior
Artul.ai Research LibraryStudy No. 69Management BehaviorUpdated 2026-08-28

Half of 165,182 earnings calls score low on evasion - and they raise guidance more

By Artul.ai Research Group · n = 86,188 earnings calls · First published 2026-08-28
Abstract

This study examines 86,188 earnings calls scoring 2 or lower on a 0-9 evasion meter, drawn from a corpus of 165,182 calls spanning 1990-2026, meaning 52.2% of all calls fall into this low-evasion group. Compared with the rest of the corpus, these calls show higher candor (7.12 vs 6.86), higher specificity (7.85 vs 7.56), lower stress (2.02 vs 2.43), and lower evasion (1.79 vs 2.70). They also raise guidance more often (24.1% vs 21.1%) and withdraw it less often (1.7% vs 2.7%). Among 10,966 calls with return data, the median follow-up return was -6.1% versus -7.2% for the 22,449-call baseline, with 40.6% beating versus 39.5%. These are descriptive differences, not causal or predictive claims.

Key findings
  • 52.2% of the 165,182-call corpus (86,188 calls) scores 2 or lower on the 0-9 evasion meter.
  • Low-evasion calls show candor of 7.12 versus 6.86 and evasion of 1.79 versus 2.70 for the rest of the corpus.
  • Guidance was raised on 24.1% of low-evasion calls versus 21.1% of others, and withdrawn on 1.7% versus 2.7%.
  • Among 10,966 low-evasion calls with returns, the median return was -6.1% versus -7.2% for the 22,449-call baseline, with 40.6% beating versus 39.5%.

1Introduction

Evasion scoring gives analysts a way to quantify how directly management answers questions, but the behavior of highly forthcoming calls has rarely been profiled at scale. If roughly half of all earnings calls are candid by this measure, the group is large enough to shape how the whole corpus reads: its tone norms, its guidance habits, and its market context. For anyone who screens calls for signals of transparency, knowing what a low-evasion call actually looks like - and how it differs from everything else - is a useful baseline. This study profiles 86,188 calls scoring 2 or lower on the 0-9 evasion meter against the remaining corpus, covering language traits, guidance actions, flagged behaviors, and follow-up returns.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 2 or lower on the 0–9 evasion meter (n = 86,188; 52.2% of the reference set, 95% Wilson interval 51.9%–52.4%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Low-evasion calls differ from the rest of the corpus on every profiled trait: candor 7.12 vs 6.86, specificity 7.85 vs 7.56, confidence 7.36 vs 7.21, with lower stress (2.02 vs 2.43), lower evasion (1.79 vs 2.70), and slightly lower promotion (4.89 vs 5.05). Guidance actions skew constructive: raised on 24.1% vs 21.1% of other calls, withdrawn on 1.7% vs 2.7%. The two flagged behaviors are 'The Question Left Hanging' (55% vs 26.5%) and 'Scale-Dependent Advantage Claims' (60% vs 6.7%). The annual share of low-evasion calls rose from 48.0% in 2015 to 54.7% in 2024 and 58.0% in 2025 (partial year). Median follow-up returns were -6.1% vs -7.2% for the 22,449-call baseline, with 40.6% beating vs 39.5%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor7.126.86+0.26
Evasion1.792.70-0.91
Specificity7.857.56+0.29
Stress2.022.43-0.41
Promotion4.895.05-0.16
Confidence7.367.21+0.15
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised24.1%21.1%
Maintained50.1%48.8%
Lowered10.4%11.6%
Withdrawn1.7%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
The Question Left Hanging0.55×26.5%48.0%
Scale-Dependent Advantage Claims0.60×6.7%11.1%
201548.03%
201649.35%
201750.49%
201850.82%
201950.13%
202050.60%
202153.90%
202253.91%
202354.34%
202454.65%
202557.95%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.1%-7.2%
Interquartile range-23.3% to +11.6%
Share beating SPY40.6% (95% CI 40%–42%)39.5%
Observations10,96622,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SBFGQ2 20252025-07-25A
USCBQ2 20252025-07-25B+
NWGQ2 20252025-07-25B+
FFICQ2 20252025-07-25B+
GBCIQ2 20252025-07-25A
MOG.AQ3 20252025-07-25B+
FLGQ2 20252025-07-25B
FRSTQ2 20252025-07-25A

4Discussion

A careful reader should conclude that calls scoring 2 or lower on the evasion meter are, by construction and by the profiled traits, more direct: higher candor and specificity, lower stress, and modestly different guidance behavior. What should not be concluded is that candor causes better outcomes or that these calls are safer investments. The return gap is small in median terms, and the beat rates differ by about one point. The over-represented flags are interesting but descriptive. No claim here supports prediction, timing, or a trading edge; the study describes what low-evasion calls look like, not what they will do.

5Limitations

The evasion score and all profiled traits are AI-read fields, which are noisy and may encode model artifacts rather than pure human behavior. The returns comparison uses 22,449 calls, a subset skewed toward liquid names, so the -6.1% versus -7.2% medians may not generalize. Our own forward tests falsified directional prediction from these signals, so nothing here should be read as an edge. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of call-based signals. All differences reported are descriptive associations within this dataset and period only. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Half of 165,182 earnings calls score low on evasion - and they raise guidance more.” Artul.ai Earnings-Call Research Library, Study No. 69. https://artul.ai/research/management-that-doesn-t-dodge

Related studies

97.97% of 161,829 earnings calls score high on speci96.2% of earnings calls score candid - yet median 1-95.5% of earnings calls since 1990 clear the 6-point39% of earnings calls trip the complexity meter - an37.3% of earnings calls score as promotional - and 2Only 11.3% of earnings calls score high on uncertain
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.