Research › Call Signals
Artul.ai Research LibraryStudy No. 85Call SignalsUpdated 2026-08-28

The Question Left Hanging: 79,206 Calls That End on an Unresolved Note

By Artul.ai Research Group · n = 79,206 earnings calls · First published 2026-08-28
Abstract

We study 79,206 earnings calls (48.0% of a 165,182-call corpus spanning 1990-2026) where the model answered YES to the battery item 'The Question Left Hanging' — calls that end with a live, unresolved issue. Compared with the base, these calls show lower candor (6.58 vs 6.86), higher evasion (3.27 vs 2.70), higher stress (2.99 vs 2.43), and lower confidence (6.88 vs 7.21). Analysts underwrite their guidance less often (65% of base lift) and such calls resolve doubts far less (0.72x lift). Post-call returns skew slightly worse (median -0.105 vs -0.072), and the share of such calls fell from 52.51% in 2015 to 45.57% in 2024.

Key findings
  • Calls with a question left hanging show higher evasion (3.27 vs 2.70) and higher stress (2.99 vs 2.43) than the corpus base.
  • Guidance is lowered on 15.4% of these calls versus 11.6% of base calls, and withdrawn on 4.2% versus 2.7%.
  • These calls are under-represented among 'Skeptic Reassured' (0.54x lift) and 'Calls That Resolve Doubts' (0.72x lift) labels.
  • Post-call returns have a median of -0.105 versus -0.072 for the 22,449-call base sample, with 36.7% beating versus 39.5%.

1Introduction

Anyone who listens to earnings calls knows the moment: the operator's queue empties, yet something raised earlier was never closed out. An analyst's sharp question gets a graceful non-answer, and the call simply ends. These hanging questions are a signature of unfinished business between management and the street, and they may accompany softer language, hedged guidance, and less reassured skeptics. Using Artul.ai's battery item 'The Question Left Hanging', this study examines 79,206 calls answered YES — 47.95% of a 165,182-call corpus — profiling their language, guidance actions, analyst labels, and post-call return distributions against the rest of the corpus.

2Data & methodology

The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "The Question Left Hanging" (n = 79,206; 48.0% of the reference set, 95% Wilson interval 47.7%–48.2%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The behavioral profile points one direction: less candor (6.58 vs 6.86), less specificity (7.19 vs 7.56), less confidence (6.88 vs 7.21), and more evasion (3.27 vs 2.70) and stress (2.99 vs 2.43). Guidance skews defensive: lowered on 15.4% of these calls vs 11.6% of base, withdrawn on 4.2% vs 2.7%, while maintained guidance is rarer (42.8% vs 48.8%). Analyst labels show the largest under-representation on 'Skeptic Reassured' (0.54x) and 'Guidance Worth Underwriting' (0.65x); the top over-represented label is 'Scale-Dependent Advantage Claims' (2.02x). The trend drifts down from 52.51% in 2015 to 45.57% in 2024, and returns skew modestly weaker (median -0.105 vs -0.072).

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.586.86-0.28
Evasion3.272.70+0.58
Specificity7.197.56-0.37
Stress2.992.43+0.56
Promotion5.215.05+0.16
Confidence6.887.21-0.33
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised12.9%21.1%
Maintained42.8%48.8%
Lowered15.4%11.6%
Withdrawn4.2%2.7%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Scale-Dependent Advantage Claims2.02×22.3%11.1%
Results Worse Than Direction1.29×65.8%51.1%
Underused Fixed Costs1.25×52.1%41.6%
Skeptic Reassured0.54×35.9%66.4%
Guidance Worth Underwriting0.65×46.5%71.5%
Calls That Resolve Doubts0.72×57.5%79.5%
201552.51%
201650.53%
201748.82%
201847.55%
201949.72%
202051.49%
202144.47%
202246.93%
202346.22%
202445.57%
202547.62%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-10.5%-7.2%
Interquartile range-31.9% to +11.7%
Share beating SPY36.7% (95% CI 36%–38%)39.5%
Observations8,04922,449
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
DOCQ2 20252025-07-25C
HCAQ2 20252025-07-25C
CNCQ2 20252025-07-25F
ULHQ2 20252025-07-25C+
DBOEYQ2 20252025-07-25B+
TBBKQ2 20252025-07-25C
LBTSFQ2 20252025-07-25C+
TRATFQ2 20252025-07-25F

4Discussion

A careful reader should treat these calls as a descriptive cluster, not a signal. Hanging questions co-occur with more evasive, more stressed language and weaker guidance posture, but co-occurrence is not causation — difficult quarters may produce both the hanging question and the softer numbers. The return gap is small, distributions overlap heavily, and no claim is made that the label predicts anything. What the data supports is a characterization: calls that end unresolved sound different, and the model's labels agree with that impression.

5Limitations

Battery fields are AI-read and noisy; a YES on 'The Question Left Hanging' reflects model judgment, not ground truth. The returns sample covers 8,049 of these calls within a 22,449-call base skewed toward liquid names, so medians may not generalize. Our own forward tests falsified directional prediction from these labels, and LLMs partially remember famous stocks' histories, contaminating any backtest. The trend mixes corpora composition across years. All findings are descriptive. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Cite this study Artul.ai Research Group (2026). “The Question Left Hanging: 79,206 Calls That End on an Unresolved Note.” Artul.ai Earnings-Call Research Library, Study No. 85. https://artul.ai/research/the-question-left-hanging-earnings-calls

Related studies

Ammunition, Not Smoke: Inside the Calls Where CriticKnow What You Know: Calls Where Confidence Matched tThe Question Left Hanging, Fewer and Farther BetweenThe Guidance Was Fine All Along: What Model-EndorsedTalking the Skeptic Down: 66% of Calls Leave the DouThe Admission Is Not the Confession: Calls Flagged a
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.