Home › Research › Earnings Call Grades

How we grade 165,000 earnings calls — and why we don't publish buy/sell

Everything on our stock pages comes from one experiment: an AI read every earnings-call transcript in our corpus — 165,000 calls across 9,117 companies — and answered the same 37 questions about each one. This page explains exactly what we measured, what predicted the future, and what didn't.

The 37 questions

For every call, the model records seven business verdicts (guidance action, demand, margins, backlog, pricing, headcount, capex), eight 0–9 behavioral readings (candor, specificity, evasion, confidence, uncertainty, promotion, stress, complexity), analyst counts, and twenty yes/no judgments — from "Would a skeptical short-seller find this call reassuring?" to "Does the CFO speak more than the CEO?". The identical battery on every call is what makes companies comparable across two decades.

The Call Grade (A–F)

The grade is a transparent composite of call quality: candor + specificity + (9 − evasion) + (9 − stress), plus bonuses when the call resolves more doubts than it creates, confidence is proportionate to evidence, a skeptic would be reassured, and guidance was raised — with penalties for critic ammunition, lowered or withdrawn guidance. Letters are fixed percentile cuts over the full corpus: A is roughly the top 7% of all calls ever, F the bottom 10%. A grade describes the quality of the conversation, not the stock.

The Expected Move (±%)

The one predictive signal that survived out-of-sample testing is about magnitude, not direction: a stock's immediate reaction to its call, combined with its pre-call volatility, predicts how far it is likely to travel over the next ~45 trading days — but not which way. We publish that estimate as "expected move ±X%". It is a volatility forecast, useful for position sizing and options context, and it is the strongest claim our data honestly supports.

What failed — and why we tell you

We tested more than 1,600 directional hypotheses: split-half validation, shuffled-label placebo runs, temporal gates across eras, pre-registration. Our best candidate — a "clean call plus growing customer prepayments" pattern we call C1 — beat the market in six separate backtest years and passed every statistical control. Then we froze a real forward watchlist and waited. It returned −12.4% median versus SPY. The backtests were also contaminated in a subtler way: large language models can partially remember what happened to famous stocks, so historical "predictions" about well-known names score far better than they deserve. That discovery is why no honest AI backtest on public companies should be taken at face value — ours included.

This is why our pages carry grades and expected moves instead of buy/sell calls. When a signal genuinely survives a forward test, we'll publish it — with the receipts.

Explore the data

All 9,117 companies A–Z All 30 screeners Guidance withdrawn Evasive management Founder-led The C1 pattern
Not investment advice. Artul.ai publishes AI-generated earnings-call quality grades and expected-volatility estimates — never buy or sell recommendations. We tested over 1,600 predictive hypotheses against 165,000 transcripts; the honest result, including what failed, is documented in our methodology.