Research › Hypotheses Tested
Artul.ai Research LibraryStudy No. 57Hypotheses TestedUpdated 2026-08-28

Getting Better Mid-Sentence: 106 Calls That Warmed Up as They Went

By Artul.ai Research Group · n = 106 earnings calls · First published 2026-08-28
Abstract

This study examines 106 earnings calls (21.8% of a 487-call corpus, 2015-2024) whose transcripts were judged YES on the hypothesis 'Still getting better as they speak' — management teams that seemed to hit their stride while the call was still underway. These speakers scored higher than the baseline on confidence (7.77 vs 7.33) and specificity (7.80 vs 7.65), and lower on stress (2.01 vs 2.38) and evasion (2.54 vs 2.69). They raised guidance far more often (35.8% vs 21.8%) and lowered it less (8.5% vs 13.6%). Overweighted research themes included 'Volume About to Step Up' (lift 1.51) and 'Pricing Recovering' (1.34).

Key findings
  • Of 487 calls in the 2015-2024 corpus, 106 (21.8%, 95% CI 18.3%-25.6%) were judged to be improving as the speakers talked.
  • These calls showed higher confidence (7.77 vs 7.33) and specificity (7.80 vs 7.65), with lower stress (2.01 vs 2.38) and evasion (2.54 vs 2.69) than the baseline.
  • Guidance was raised on 35.8% of these calls versus 21.8% of the baseline, and lowered on 8.5% versus 13.6%.
  • In the 34-call returns sample, 47.1% beat the market versus 37.8% of the 193-call baseline, with a median return of -1.9% against -10.5%.

1Introduction

Anyone who reads transcripts for a living knows the feeling of a call that starts stiff and slowly loosens up — the CEO finds the thread, the answers get sharper, the numbers start coming unprompted. That arc is easy to sense and hard to measure. If 'getting better as they speak' marks a management team with genuine command of its story, it might line up with more candid language, more specific detail, and more aggressive guidance posture. This study profiles the 106 calls in our 2015-2024 corpus that answered YES to that hypothesis, comparing their language scores, guidance actions, and research-theme weights against the other 487-call baseline.

2Data & methodology

The corpus comprises 487 earnings-call transcripts published between 2015 and 2024, each scored independently by a large language model on an identical 37-field battery: seven categorical business verdicts, eight 0–9 behavioral meters, and twenty yes/no judgments. The study group is defined as calls that answered YES to the research hypothesis "Still getting better as they speak" (n = 106; 21.8% of the reference set, 95% Wilson interval 18.3%–25.6%). Baseline figures use the set of calls on which this question was tested. Market outcomes join a fixed sample of 22,449 calls with twelve-month total returns in excess of SPY, measured from the first close after each call; this sample skews toward liquid U.S. names and is reported as descriptive history only.

3Results

The clearest signal is in guidance posture: 35.8% of these calls raised guidance versus 21.8% of the baseline, while only 8.5% lowered it versus 13.6%. Language scores point the same direction — confidence (7.77 vs 7.33), specificity (7.80 vs 7.65), and promotion (5.44 vs 5.09) all run above baseline, with stress (2.01 vs 2.38) and evasion (2.54 vs 2.69) below it. Overweighted themes include 'Volume About to Step Up' (lift 1.51) and 'Pricing Recovering' (1.34); underweighted are 'Results Worse Than Direction' (0.66) and 'When the CFO Dominates' (0.70). The annual share peaked at 0.19 in 2022 and was 0.06 in 2024.

Table 1. Mean behavioral scores (0–9 scale), study group versus baseline
MeterStudy groupBaselineΔ
Candor6.946.92+0.02
Evasion2.542.69-0.16
Specificity7.807.65+0.16
Stress2.012.38-0.37
Promotion5.445.09+0.36
Confidence7.777.33+0.45
Table 2. Guidance actions, study group versus baseline
ActionStudy groupBaseline
Raised35.8%21.8%
Maintained47.2%52.8%
Lowered8.5%13.6%
Withdrawn0.9%1.0%
Table 3. Co-occurring battery signals ranked by lift (group prevalence ÷ baseline prevalence)
SignalLiftIn groupBaseline
Volume About to Step Up1.51×37.7%25.1%
Pricing Recovering1.34×22.6%16.8%
Founder-Led Companies1.26×26.4%20.9%
Results Worse Than Direction0.66×31.1%46.8%
When the CFO Dominates0.70×10.4%14.8%
20150.09%
20160.06%
20170.05%
20180.05%
20190.00%
20200.00%
20210.08%
20220.19%
20230.09%
20240.06%
20250.00%
Figure 1. Share of all analyzed calls matching the study definition, by year.
Table 4. Twelve-month excess total returns versus SPY (descriptive history, not a signal)
StatisticStudy groupReturns sample
Median excess return-1.9%-10.5%
Interquartile range-26.7% to +12.5%
Share beating SPY47.1% (95% CI 31%–63%)37.8%
Observations34193
Table 5. Most recent calls matching the study definition
TickerQuarterCall dateCall grade
WRBYQ1 20242024-05-09A
AZEKQ2 20242024-05-08B+
PPCQ1 20242024-05-03A
CTRAQ1 20242024-05-03A
CWQ1 20242024-05-02B+
ROCKQ1 20242024-05-01B+
PDSQ1 20242024-04-25B
ASMQ4 20232024-03-21C+

4Discussion

A careful reader should conclude that calls judged to improve mid-stream also tend to carry confident, specific, low-stress language and a more favorable guidance mix — these are descriptions of the same calls, measured at the same time. What should not be concluded is that this vocal arc causes good outcomes, predicts returns, or can be traded. The 47.1% beat rate in the 34-call returns sample sits within a wide confidence interval (31.5%-63.3%) that overlaps the baseline's 37.8%. The pattern is a coherent profile, not a proven edge.

5Limitations

The hypothesis labels and language scores are produced by AI reading transcripts and are noisy; a YES may reflect topic, formatting, or drafting quirks as much as speaker behavior. The returns sample covers 34 of these calls (against 193 baseline) drawn from a universe of 22,449 calls skewed toward liquid names, so it is small and not representative. Our own forward tests falsified directional prediction from these signals. Additionally, LLMs partially remember famous stocks' histories, contaminating any backtest built on their judgments. See the full methodology, including the C1 pattern’s forward-test failure and the LLM-memorization finding.

Companion page: every company matching this hypothesis is listed at the question’s own page.

Cite this study Artul.ai Research Group (2026). “Getting Better Mid-Sentence: 106 Calls That Warmed Up as They Went.” Artul.ai Earnings-Call Research Library, Study No. 57. https://artul.ai/research/hypothesis-still-getting-better-as-they-speak

Related studies

The Numbers Are Fine; Everything Else Is Pending: FlRoom to Run and Nowhere to Hide: Calls With UncontesLight at the End of the Tunnel Is Often a Train: RecDon't Stick to the Script: Earnings Calls Where AnswHave Your Cake and Expand the Base Too: Repeat GrowtPressed Harder, Answered Straight: A Profile of 183
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.