Research › Call Signals
Artul.ai Research LibraryStudy No. 14Call SignalsUpdated 2026-08-28

Calls that resolve doubts: 79.5% of 165,182 earnings calls, and 40.3% raised guidance

By Artul.ai Research Group · n = 131,335 earnings calls · First published 2026-08-28
Abstract

This study examines 165,182 earnings calls from 1990 to 2026 and isolates the 131,335 calls (79.5%, 95% CI 79.3%–79.7%) where the model answered YES to 'Calls That Resolve Doubts.' Compared with other calls, these show higher candor (6.99 vs 6.86), specificity (7.75 vs 7.56), and confidence (7.42 vs 7.21), with lower evasion (2.49 vs 2.70) and stress (2.12 vs 2.43). Guidance behavior differs too: 24.9% raised guidance versus 21.1% elsewhere. The over-represented topic 'The Question Left Hanging' appears at 0.72% versus 0.35% (lift 2.08). A median one-day return of -0.064% was slightly above the -0.072% baseline, and 40.3% beat, versus 39.5% baseline.

Key findings
  • 79.5% of 165,182 earnings calls (n=131,335) were classified as resolving doubts, with a 95% CI of 79.3% to 79.7%.
  • Doubt-resolving calls show higher candor (6.99 vs 6.86), specificity (7.75 vs 7.56), and confidence (7.42 vs 7.21), and lower evasion (2.49 vs 2.70) and stress (2.12 vs 2.43).
  • 24.9% of doubt-resolving calls raised guidance versus 21.1% of other calls, while 8.9% lowered guidance versus 11.6% elsewhere.
  • In a 19,998-call returns subsample, the median one-day return was -0.064% versus -0.072% baseline, with 40.3% beating versus 39.5% baseline.

1Introduction

Earnings calls are the primary venue where management addresses analyst skepticism, and whether a call actually resolves doubts is a signal many listeners track informally. With 165,182 calls scored by a language model across 1990–2026, we can quantify how common this quality is and what linguistic and disclosure patterns accompany it. If doubt-resolving calls differ systematically in candor, evasion, and guidance behavior, that tells us how disclosure quality clusters. This study examines the 131,335 calls the model labeled YES to 'Calls That Resolve Doubts,' their language profile, guidance actions, topic over/under-representation, yearly trend, and one-day returns.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Calls That Resolve Doubts" (n = 131,335; 79.5% of the reference set, 95% Wilson interval 79.3%–79.7%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Doubt-resolving calls (79.5% of the corpus) skew toward clarity: candor 6.99 vs 6.86, specificity 7.75 vs 7.56, confidence 7.42 vs 7.21, with evasion at 2.49 vs 2.70 and stress at 2.12 vs 2.43. Guidance actions differ modestly: 24.9% raised guidance (vs 21.1%) and 52.3% maintained it (vs 48.8%), while only 8.9% lowered (vs 11.6%) and 2.2% withdrew (vs 2.7%). 'The Question Left Hanging' is over-represented at 0.72% vs 0.35% (lift 2.08), and 'Scale-Dependent Advantage Claims' at 0.33% vs 0.11% (lift 3.00). The share dipped to 74.25% in 2025 after peaking at 83.34% in 2020. Median one-day returns were -0.064% vs -0.072% baseline, with 40.3% beating vs 39.5%.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.996.86+0.13
Evasion2.492.70-0.21
Specificity7.757.56+0.19
Stress2.122.43-0.31
Promotion5.015.05-0.04
Confidence7.427.21+0.21
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised24.9%21.1%
Maintained52.3%48.8%
Lowered8.9%11.6%
Withdrawn2.2%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims0.33×3.6%11.1%
The Question Left Hanging0.72×34.7%48.0%
201576.43%
201677.93%
201779.34%
201879.57%
201978.55%
202083.34%
202183.05%
202278.05%
202379.20%
202478.81%
202574.25%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.4%-7.2%
Interquartile range-24.4% to +12.2%
Share beating SPY40.3% (95% CI 40%–41%)39.5%
Observations19,99822,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SBFGQ2 20252025-07-25A
DOCQ2 20252025-07-25C
USCBQ2 20252025-07-25B+
AONQ2 20252025-07-25C
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B
FFICQ2 20252025-07-25B+
OMFQ2 20252025-07-25A

4Discussion

A careful reader should conclude that doubt-resolving calls are the modal case and that they co-occur with somewhat more candor, specificity, confidence, and guidance raises. These are associations measured on the same calls, not evidence that resolving doubts causes better outcomes or that the language profile causes guidance actions. The returns differences are small (-0.064% vs -0.072% median; 40.3% vs 39.5% beat rate) and should not be read as an exploitable pattern. The 2025 figure (74.25%) rests on only 6,012 calls and may reflect partial-year data rather than a real decline.

5Limitations

The YES/NO labels come from AI-read fields and are noisy; a model's judgment of whether a call resolves doubts may not match a human analyst's. The returns subsample covers 22,449 calls and is skewed toward liquid names, so return statistics may not generalize. Our own forward tests falsified directional prediction, so nothing here should be treated as a trading edge. Additionally, LLMs partially remember famous stocks' histories, which can contaminate any backtest of model-scored text. The 2025 trend point (74.25%, n=6,012) may reflect incomplete data. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Calls that resolve doubts: 79.5% of 165,182 earnings calls, and 40.3% raised guidance.” Artul.ai Earnings-Call Research Library, Study No. 14. https://artul.ai/research/calls-that-resolve-doubts-earnings-calls

Related studies

87.8% of earnings calls contain critic ammunition - 87.6% of calls earn a YES on proportionate confidenc71.5% of earnings calls got a yes on underwriting guWhen the skeptic buys it: 66.4% of earnings calls enHalf of earnings calls flagged for worse-than-guidedWhen a call leaves one question hanging, raised guid
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.