Research › Call Signals
Artul.ai Research LibraryStudy No. 19Call SignalsUpdated 2026-08-28

Know What You Know: Calls Where Confidence Matched the Evidence

By Artul.ai Research Group · n = 144,714 earnings calls · First published 2026-08-28
Abstract

This study examines earnings calls where a language model answered YES to the battery item 'Confidence Proportionate to Evidence'. Of 165,182 calls scored from 1990 to 2026, 144,714 (87.6%, 95% CI 87.4% to 87.8%) met that bar. Relative to the remainder, these calls show higher candor (6.97 vs 6.86), higher specificity (7.67 vs 7.56), and lower stress (2.28 vs 2.43). Guidance was raised on 22.8% of matched calls vs 21.1% of others. Among 21,399 calls with post-call returns, the median return was -0.067 vs -0.072 for the base set. One battery item, 'Scale-Dependent Advantage Claims', appeared below its expected rate.

Key findings
  • 144,714 of 165,182 calls (87.6%, CI 87.4%-87.8%) were flagged as having confidence proportionate to evidence.
  • Matched calls scored higher on candor (6.97 vs 6.86) and specificity (7.67 vs 7.56), and lower on stress (2.28 vs 2.43).
  • Guidance was raised on 22.8% of matched calls versus 21.1% of the rest; lowered guidance was slightly rarer (11.0% vs 11.6%).
  • Among calls with post-call returns, the median return was -0.067 for matched calls versus -0.072 in the 22,449-call base set.
  • 'Scale-Dependent Advantage Claims' appeared 38% below its expected rate on matched calls (observed 4.2% vs 11.1% expected).

1Introduction

Anyone who parses earnings calls knows the feeling: a management team radiating certainty that the transcript does not quite support. A model-scored battery offers a systematic way to flag the opposite case - calls where expressed confidence lines up with the evidence presented. Because 87.6% of the corpus passes this filter, the interesting question is less who passes than what distinguishes the passing calls from the rest. This study examines 144,714 such calls against the remaining 20,468, comparing language profiles, guidance actions, theme rates, and post-call returns.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "Confidence Proportionate to Evidence" (n = 144,714; 87.6% of the reference set, 95% Wilson interval 87.4%–87.8%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

Matched calls tilt toward plainer speech: candor runs 6.97 versus 6.86, specificity 7.67 versus 7.56, while promotion language (4.89 vs 5.05) and stress (2.28 vs 2.43) both run lower. Guidance actions lean modestly positive: 22.8% raised versus 21.1% elsewhere, and 50.8% maintained versus 48.8%. Only one theme deviated from its expected rate - 'Scale-Dependent Advantage Claims' at 38% below expectation (4.2% observed vs 11.1% expected). The annual share is fairly stable, from 87.14% in 2015 to 86.56% in 2025, with a 2020 peak of 90.5% and a 2022 trough of 85.86%. Median post-call returns were -0.067 for matched calls versus -0.072 for the base.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.976.86+0.11
Evasion2.602.70-0.10
Specificity7.677.56+0.11
Stress2.282.43-0.15
Promotion4.895.05-0.16
Confidence7.257.21+0.04
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised22.8%21.1%
Maintained50.8%48.8%
Lowered11.0%11.6%
Withdrawn2.6%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims0.38×4.2%11.1%
201587.14%
201687.99%
201787.67%
201888.07%
201987.87%
202090.50%
202187.56%
202285.86%
202386.76%
202486.91%
202586.56%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-6.7%-7.2%
Interquartile range-25.1% to +12.1%
Share beating SPY40.0% (95% CI 39%–41%)39.5%
Observations21,39922,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
SBFGQ2 20252025-07-25A
DOCQ2 20252025-07-25C
USCBQ2 20252025-07-25B+
HCAQ2 20252025-07-25C
AONQ2 20252025-07-25C
NWGQ2 20252025-07-25B+
BFHQ2 20252025-07-25B
FFICQ2 20252025-07-25B+

4Discussion

A careful reader should treat these results as descriptive. Calls rated as proportionate in confidence also read as more candid and specific, less promotional, and slightly less stressful - a coherent profile, but a correlation, not proof that one causes the other. The single under-weighted theme suggests such calls rarely lean on scale-dependent advantage claims. Post-call returns are statistically indistinguishable in practical terms, and none of these numbers support trading on the flag. The 2020 peak and 2022 trough are worth noting but not interpreting as market signals.

5Limitations

Battery items are model-generated and noisy; the YES label reflects one reading of one prompt. The returns sample covers only 21,399 of the flagged calls (22,449 in the base) and skews toward liquid names, so price outcomes may not generalize. Our own forward tests falsified directional prediction, and LLMs partially remember famous stocks' histories, contaminating any backtest. Corpus coverage also varies by year, with 2025 partial at 6,012 calls, so trend comparisons across periods are uneven. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “Know What You Know: Calls Where Confidence Matched the Evidence.” Artul.ai Earnings-Call Research Library, Study No. 19. https://artul.ai/research/confidence-proportionate-to-evidence-earnings-calls

Related studies

Ammunition, Not Smoke: Inside the Calls Where CriticThe Question Left Hanging, Fewer and Farther BetweenThe Guidance Was Fine All Along: What Model-EndorsedTalking the Skeptic Down: 66% of Calls Leave the DouThe Admission Is Not the Confession: Calls Flagged aThe Question Left Hanging: 79,206 Calls That End on
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.