Trust the Raise: When Skeptics Get Reassured and Guidance Goes Up
This study isolates earnings calls where Artul.ai's model answered YES to "Skeptic Reassured" and guidance was read as raised: 31,112 calls from a 165,182-call corpus spanning 1990-2026, or 18.8% of calls (95% CI 18.6%-19.0%). These calls differ sharply in tone: stress sits at 1.72 versus a 2.43 baseline (a -0.71 gap), evasion runs 2.39 versus 2.70, and confidence runs 7.85 versus 7.21. By construction, 100% raised guidance versus 21.1% in the base. On outcomes, 42.1% of the 6,378 return-sample calls beat, versus 39.5% baseline, with a median return of -5.3% versus -7.2%. "Deferred Revenue Growing" shows the strongest over-representation at 1.59x.
- The pattern covers 31,112 calls, 18.8% of the 165,182-call corpus, with a 95% confidence interval of 18.6% to 19.0%.
- Tone deltas are pronounced: stress is 1.72 versus 2.43 baseline, evasion 2.39 versus 2.70, and confidence 7.85 versus 7.21.
- "Deferred Revenue Growing" appears 1.59x more often than baseline, while "Scale-Dependent Advantage Claims" appears at just 0.08x.
- In the 6,378-call return sample, 42.1% beat versus a 39.5% base rate, and the median return is -5.26% versus -7.16% baseline.
1Introduction
Analysts often treat a raised guidance number as the headline and move on. But the tone surrounding that raise—whether the skeptics in the room came away satisfied—may tell a different, subtler story about how management presented the news. Calls where hard questions end with the questioner reassured, and where guidance is simultaneously raised, form a distinct population in the earnings-call record. Using Artul.ai's model-read fields across 165,182 calls from 1990 to 2026, this study profiles those 31,112 calls: their language patterns, the qualitative flags that cluster with them, how the population shifts year to year, and how their post-call returns compare with the broader corpus.
2Data & methodology
The corpus comprises 165,182 earnings-call transcripts published between 1990 and 2026, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls where the model answered YES to "Skeptic Reassured" AND guidance was read as raised (n = 31,112; 18.8% of the reference set, 95% Wilson interval 18.6%–19.0%). Baseline figures use all scored calls. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.
3Results
The tone profile reads as unusually calm and specific: stress is 1.72 against a 2.43 baseline, evasion 2.39 versus 2.70, while specificity (7.89 vs 7.56), candor (6.98 vs 6.86), and confidence (7.85 vs 7.21) all run higher. Flag lift is asymmetric. "Deferred Revenue Growing" runs at 1.59x baseline and "Guidance Worth Underwriting" at 1.35x, while defensive flags collapse: "Scale-Dependent Advantage Claims" at 0.08x and "Results Worse Than Direction" at 0.41x. The share of calls peaked at 28.8% in 2021 before easing to 15.19% in 2025. In the return sample, 42.1% beat versus 39.5% baseline, with a median of -5.26% versus -7.16%.
| Meter | Study group | Baseline | Δ |
|---|---|---|---|
| Candor | 6.98 | 6.86 | +0.12 |
| Evasion | 2.39 | 2.70 | -0.31 |
| Specificity | 7.89 | 7.56 | +0.34 |
| Stress | 1.72 | 2.43 | -0.71 |
| Promotion | 5.28 | 5.05 | +0.23 |
| Confidence | 7.85 | 7.21 | +0.63 |
| Action | Study group | Baseline |
|---|---|---|
| Raised | 100.0% | 21.1% |
| Maintained | 0.0% | 48.8% |
| Lowered | 0.0% | 11.6% |
| Withdrawn | 0.0% | 2.7% |
| Signal | Lift | In group | Baseline |
|---|---|---|---|
| Deferred Revenue Growing | 1.59× | 14.1% | 8.9% |
| Guidance Worth Underwriting | 1.35× | 96.6% | 71.5% |
| Early Products Growing Fast | 1.26× | 48.6% | 38.5% |
| Pricing Recovering | 1.25× | 27.0% | 21.5% |
| Scale-Dependent Advantage Claims | 0.08× | 0.8% | 11.1% |
| Results Worse Than Direction | 0.41× | 20.7% | 51.1% |
| The Question Left Hanging | 0.46× | 21.9% | 48.0% |
| The Hidden Segment | 0.66× | 13.9% | 21.1% |
| Underused Fixed Costs | 0.66× | 27.6% | 41.6% |
| Statistic | Study group | Returns sample |
|---|---|---|
| Median excess return | -5.3% | -7.2% |
| Interquartile range | -23.8% to +14.1% | — |
| Share beating SPY | 42.1% (95% CI 41%–43%) | 39.5% |
| Observations | 6,378 | 22,449 |
| Ticker | Quarter | Call date | Call grade |
|---|---|---|---|
| SBFG | Q2 2025 | 2025-07-25 | A |
| HCA | Q2 2025 | 2025-07-25 | C |
| NWG | Q2 2025 | 2025-07-25 | B+ |
| BFH | Q2 2025 | 2025-07-25 | B |
| FFIC | Q2 2025 | 2025-07-25 | B+ |
| OMF | Q2 2025 | 2025-07-25 | A |
| GBCI | Q2 2025 | 2025-07-25 | A |
| MOG.A | Q3 2025 | 2025-07-25 | B+ |
4Discussion
A careful reader should conclude only that this combination of model-read tone and raised guidance coincides with calmer language, fewer defensive flags, and a modestly better beat rate—descriptive differences, not causes. The beat-rate gap of 2.6 points is small and the median return remains negative, so nothing here suggests an actionable edge. The flag lifts describe which framings co-occur with reassured skeptics, not which framings produce good outcomes. Year-to-year swings likely reflect market regime as much as behavior.
5Limitations
The underlying fields—Skeptic Reassured, guidance direction, and the qualitative flags—are AI-read labels and carry real annotation noise. The returns sample covers 22,449 calls and is skewed toward liquid names, limiting generalizability. Artul.ai's own forward tests have falsified directional prediction on this data, so no trading implication should be drawn. Finally, LLMs partially remember famous stocks' histories, which contaminates any backtest-style comparison of returns across flagged groups. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.