The Question Left Hanging: 79,206 Calls That End on an Unresolved Note
We study 79,206 earnings calls (48.0% of a 165,182-call corpus spanning 1990-2026) where the model answered YES to the battery item 'The Question Left Hanging' — calls that end with a live, unresolved issue. Compared with the base, these calls show lower candor (6.58 vs 6.86), higher evasion (3.27 vs 2.70), higher stress (2.99 vs 2.43), and lower confidence (6.88 vs 7.21). Analysts underwrite their guidance less often (65% of base lift) and such calls resolve doubts far less (0.72x lift). Post-call returns skew slightly worse (median -0.105 vs -0.072), and the share of such calls fell from 52.51% in 2015 to 45.57% in 2024.
- Calls with a question left hanging show higher evasion (3.27 vs 2.70) and higher stress (2.99 vs 2.43) than the corpus base.
- Guidance is lowered on 15.4% of these calls versus 11.6% of base calls, and withdrawn on 4.2% versus 2.7%.
- These calls are under-represented among 'Skeptic Reassured' (0.54x lift) and 'Calls That Resolve Doubts' (0.72x lift) labels.
- Post-call returns have a median of -0.105 versus -0.072 for the 22,449-call base sample, with 36.7% beating versus 39.5%.
1Introduction
Anyone who listens to earnings calls knows the moment: the operator's queue empties, yet something raised earlier was never closed out. An analyst's sharp question gets a graceful non-answer, and the call simply ends. These hanging questions are a signature of unfinished business between management and the street, and they may accompany softer language, hedged guidance, and less reassured skeptics. Using Artul.ai's battery item 'The Question Left Hanging', this study examines 79,206 calls answered YES — 47.95% of a 165,182-call corpus — profiling their language, guidance actions, analyst labels, and post-call return distributions against the rest of the corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to the battery item "The Question Left Hanging" (n = 79,206; 48.0% of the reference set, 95% Wilson interval 47.7%–48.2%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The behavioral profile points one direction: less candor (6.58 vs 6.86), less specificity (7.19 vs 7.56), less confidence (6.88 vs 7.21), and more evasion (3.27 vs 2.70) and stress (2.99 vs 2.43). Guidance skews defensive: lowered on 15.4% of these calls vs 11.6% of base, withdrawn on 4.2% vs 2.7%, while maintained guidance is rarer (42.8% vs 48.8%). Analyst labels show the largest under-representation on 'Skeptic Reassured' (0.54x) and 'Guidance Worth Underwriting' (0.65x); the top over-represented label is 'Scale-Dependent Advantage Claims' (2.02x). The trend drifts down from 52.51% in 2015 to 45.57% in 2024, and returns skew modestly weaker (median -0.105 vs -0.072).
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.58 | 6.86 | -0.28 |
| Evasion | 3.27 | 2.70 | +0.58 |
| Specificity | 7.19 | 7.56 | -0.37 |
| Stress | 2.99 | 2.43 | +0.56 |
| Promotion | 5.21 | 5.05 | +0.16 |
| Confidence | 6.88 | 7.21 | -0.33 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 12.9% | 21.1% |
| Maintained | 42.8% | 48.8% |
| Lowered | 15.4% | 11.6% |
| Withdrawn | 4.2% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Scale-Dependent Advantage Claims | 2.02× | 22.3% | 11.1% |
| Results Worse Than Direction | 1.29× | 65.8% | 51.1% |
| Underused Fixed Costs | 1.25× | 52.1% | 41.6% |
| Skeptic Reassured | 0.54× | 35.9% | 66.4% |
| Guidance Worth Underwriting | 0.65× | 46.5% | 71.5% |
| Calls That Resolve Doubts | 0.72× | 57.5% | 79.5% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -10.5% | -7.2% |
| Interquartile range | -31.9% to +11.7% | — |
| Share beating SPY | 36.7% (95% CI 36%–38%) | 39.5% |
| Observations | 8,049 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| DOC | Q2 2025 | 2025-07-25 | C |
| HCA | Q2 2025 | 2025-07-25 | C |
| CNC | Q2 2025 | 2025-07-25 | F |
| ULH | Q2 2025 | 2025-07-25 | C+ |
| DBOEY | Q2 2025 | 2025-07-25 | B+ |
| TBBK | Q2 2025 | 2025-07-25 | C |
| LBTSF | Q2 2025 | 2025-07-25 | C+ |
| TRATF | Q2 2025 | 2025-07-25 | F |
4Discussion
A careful reader should treat these calls as a descriptive cluster, not a signal. Hanging questions co-occur with more evasive, more stressed language and weaker guidance posture, but co-occurrence is not causation — difficult quarters may produce both the hanging question and the softer numbers. The return gap is small, distributions overlap heavily, and no claim is made that the label predicts anything. What the data supports is a characterization: calls that end unresolved sound different, and the model's labels agree with that impression.
5Limitations
Battery fields are AI-read and noisy; a YES on 'The Question Left Hanging' reflects model judgment, not ground truth. The returns sample covers 8,049 of these calls within a 22,449-call base skewed toward liquid names, so medians may not generalize. Our own forward tests falsified directional prediction from these labels, and LLMs partially remember famous stocks' histories, contaminating any backtest. The trend mixes corpora composition across years. All findings are descriptive. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.