Research › Management Behavior
Artul.ai Research LibraryStudy No. 88Management BehaviorUpdated 2026-08-28

96.2% of earnings calls score candid - yet median 1-day return is -0.07%

By Artul.ai Research Group · n = 158,979 earnings calls · First published 2026-08-28
Abstract

This study profiles earnings calls that score 6 or higher on a 0-9 candor meter, drawing on 158,979 qualifying calls out of a 165,182-call corpus spanning 1990 to 2026. High-candor calls show modestly elevated language profiles: candor 6.95 vs 6.86 baseline, specificity 7.62 vs 7.56, and confidence 7.22 vs 7.21, with slightly lower evasion (2.66 vs 2.70) and stress (2.40 vs 2.43). Guidance behavior differs little: 21.4% of qualifying calls raised guidance vs 21.1% baseline. Among 22,192 calls with return data, the median one-day return was -0.07%, essentially matching the -0.07% baseline, and 39.6% beat, nearly identical to the 39.5% base rate. The candor meter's reach is wide, but its market footprint appears negligible.

Key findings
  • 158,979 of 165,182 calls (96.2%, 95% CI 96.2%-96.3%) score 6 or higher on the 0-9 candor meter.
  • High-candor calls average specificity of 7.62 versus a 7.56 baseline, the largest profile gap in the study (+0.07).
  • 21.4% of high-candor calls raised guidance versus 21.1% of all calls, with 49.4% maintaining versus 48.8% overall.
  • Among 22,192 high-candor calls with returns, the median one-day return was -0.07% and 39.6% beat, versus 39.5% for the full 22,449-call sample.

1Introduction

Candor scores promise a simple read on management trustworthiness, so it matters how many calls actually clear a high bar and what, if anything, that signals. If nearly every call scored as candid, the meter would say little; if high-candor calls clustered with guidance raises or outsized returns, the score could carry information. Prior work on earnings-call language shows tone metrics often look distinctive in text but faint in market outcomes. This study examines calls scoring 6 or higher on the 0-9 candor meter across 165,182 calls from 1990 to 2026, comparing their language profiles, guidance actions, and short-horizon returns against the full corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls scoring 6 or higher on the 0–9 candor meter (n = 158,979; 96.2% of the reference set, 95% Wilson interval 96.2%–96.3%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The qualifying share is 96.2% of all calls, so the 6+ threshold separates few calls from the corpus. Language deltas are correspondingly small: candor 6.95 vs 6.86, specificity 7.62 vs 7.56, confidence 7.22 vs 7.21, with evasion down 0.04 and stress down 0.03. Guidance actions track the baseline closely: 21.4% raised, 49.4% maintained, 11.9% lowered, and 2.7% withdrew, against 21.1%, 48.8%, 11.6%, and 2.7% corpus-wide. The annual share of qualifying calls stayed between 95.5% and 97.5% from 2015 through 2024, with 2025 at 89.5% on 6,012 calls. Returns show no separation: median -0.07% vs -0.07% baseline, and beat rates of 39.6% vs 39.5%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.956.86+0.09
Evasion2.662.70-0.04
Specificity7.627.56+0.07
Stress2.402.43-0.03
Promotion5.015.05-0.04
Confidence7.227.21+0.01
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised21.4%21.1%
Maintained49.4%48.8%
Lowered11.9%11.6%
Withdrawn2.7%2.7%
201596.44%
201696.69%
201796.76%
201896.69%
201996.83%
202097.50%
202196.06%
202296.21%
202396.44%
202495.49%
202589.45%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 3. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-7.1%-7.2%
Interquartile range-25.9% to +11.9%
Share beating SPY39.6% (95% CI 39%–40%)39.5%
Observations22,19222,449
Table 4. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SBFGQ2 20252025-07-25A
DOCQ2 20252025-07-25C
USCBQ2 20252025-07-25B+
HCAQ2 20252025-07-25C
AONQ2 20252025-07-25C
CNCQ2 20252025-07-25F
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B

4Discussion

A careful reader should conclude that the 6+ candor threshold describes most calls and that qualifying calls differ only marginally in language, guidance behavior, and next-day returns. The near-identical beat rates (39.6% vs 39.5%) and medians (-0.07% vs -0.07%) suggest the threshold, as set, does not partition calls into distinct groups. What should not be concluded: that candor causes returns, that the meter predicts market reactions, or that the 2025 dip to 89.5% signals a lasting regime change, since that year covers only 6,012 calls. Differences of 0.01-0.09 on profile scales are within the range one would expect from threshold selection alone.

5Limitations

Candor, evasion, and specificity are AI-read fields and inherently noisy; small deltas may reflect scoring error. The returns sample covers 22,192 of 158,979 qualifying calls (22,449 baseline) and is skewed toward liquid, larger names. Our own forward tests falsified directional prediction from these features, and LLMs partially remember famous stocks' histories, contaminating any backtest. The 2025 figure rests on a partial year of 6,012 calls. All comparisons are descriptive; no causal or predictive claim is supported by these statistics. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “96.2% of earnings calls score candid - yet median 1-day return is -0.07%.” Artul.ai Earnings-Call Research Library, Study No. 88. https://artul.ai/research/unusually-candid-management

Related studies

97.97% of 161,829 earnings calls score high on speci95.5% of earnings calls since 1990 clear the 6-pointHalf of 165,182 earnings calls score low on evasion 39% of earnings calls trip the complexity meter - an37.3% of earnings calls score as promotional - and 2Only 11.3% of earnings calls score high on uncertain
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.