Everything on our stock pages comes from one experiment: an AI read every earnings-call transcript in our corpus — 165,000 calls across 9,117 companies — and answered the same 37 questions about each one. This page explains exactly what we measured, what predicted the future, and what didn't.
For every call, the model records seven business verdicts (guidance action, demand, margins, backlog, pricing, headcount, capex), eight 0–9 behavioral readings (candor, specificity, evasion, confidence, uncertainty, promotion, stress, complexity), analyst counts, and twenty yes/no judgments — from "Would a skeptical short-seller find this call reassuring?" to "Does the CFO speak more than the CEO?". The identical battery on every call is what makes companies comparable across two decades.
The grade is a transparent composite of call quality: candor + specificity + (9 − evasion) + (9 − stress), plus bonuses when the call resolves more doubts than it creates, confidence is proportionate to evidence, a skeptic would be reassured, and guidance was raised — with penalties for critic ammunition, lowered or withdrawn guidance. Letters are fixed percentile cuts over the full corpus: A is roughly the top 7% of all calls ever, F the bottom 10%. A grade describes the quality of the conversation, not the stock.
The one predictive signal that survived out-of-sample testing is about magnitude, not direction: a stock's immediate reaction to its call, combined with its pre-call volatility, predicts how far it is likely to travel over the next ~45 trading days — but not which way. We publish that estimate as "expected move ±X%". It is a volatility forecast, useful for position sizing and options context, and it is the strongest claim our data honestly supports.
We tested more than 1,600 directional hypotheses: split-half validation, shuffled-label placebo runs, temporal gates across eras, pre-registration. Our best candidate — a "clean call plus growing customer prepayments" pattern we call C1 — beat the market in six separate backtest years and passed every statistical control. Then we froze a real forward watchlist and waited. It returned −12.4% median versus SPY. The backtests were also contaminated in a subtler way: large language models can partially remember what happened to famous stocks, so historical "predictions" about well-known names score far better than they deserve. That discovery is why no honest AI backtest on public companies should be taken at face value — ours included.
This is why our pages carry grades and expected moves instead of buy/sell calls. When a signal genuinely survives a forward test, we'll publish it — with the receipts.