Methodology 4.0 · adopted 2026-08-02The ruler, visible.
Every score on this site is computed from grades that carry a verbatim source sentence. No quote, no verdict. This page states the whole method, the declared eras behind it, and the parts we got wrong on the way here.
The instrument: four agents, each blind to the next
An extraction agent reads every earnings call and records each forward-looking claim with its verbatim sentence and the test it commits to. A reading agent later takes each due claim and the company's subsequent transcripts, and must return the actual outcome with the verbatim resolving sentence. A grading agent receives only the committed sentence and the resolving sentence, and states the verdict and the signed deviation. An independent verifier re-derives what each open claim commits to before it appears in any client report. Language is judged by models with citable evidence or abstention; code decides only equality, arithmetic, and dates.
The ruler: band, decay, materiality
The official band is 5 percent around the committed level: inside it a claim is MET, outside it is BEAT or MISSED by the sign of the stated deviation. Severity decays exponentially beyond the band, so a miss by 20 costs far more than a miss by 6. Claims are weighted by materiality tier, core guidance at 1.0, operational commitments at 0.6, peripheral items at 0.3, judged per claim by a model, never by keyword. Qualitative commitments grade on their own stated bars and count at the label's value when no honest percentage exists. Directional reversals cap at 200. The transparency slider on the rankings recomputes labels from the stated deviations at any band you choose; the official column never moves.
Revisions, bias, and the sign convention
A commitment resolved by the company revising it grades at the moved level and displays as REVISED UP or REVISED DOWN, at identical weight to any other grade: walking a number down is not a free action. Bias is the median signed deviation across an executive's numeric grades; positive means outcomes land favorable to guidance. Favorability owns the sign everywhere, including costs: incurring less of a guided cost is favorable, whatever the raw arithmetic says.
Abstention is honest non-publication
When the record cannot support a verdict, the grader abstains, and the abstention rate is disclosed per company rather than hidden. Roughly a quarter of resolutions across the full corpus end in abstention; that floor is a property of how executives speak, and publishing a guess instead would be the actual failure.
rho 0.371
95 percent confidence interval [0.243, 0.496] · pre-registered publication gate 0.30 · temporal split
What this means. We split every executive’s graded claims in half by time and ask whether accuracy on the earlier half predicts accuracy on the later half. A rho of 1.0 would mean rankings persist perfectly; a rho of 0 would mean the leaderboard measures luck. The interval sitting clearly above zero means executive accuracy is a persistent trait: the same people keep being right, and the same people keep being wrong. The 2026-08-01 check on the retired verification layer, same scoring formula, measured 0.410 [0.280, 0.525]; two verification approaches with overlapping intervals reaching one conclusion is checkable robustness, not a coincidence. Persistence and its interval are recomputed live on the rankings page as new grades land.
Eras: declared changes, never quiet retuning
Methodology 4.0, adopted 2026-08-02, is a ground-up rebuild. The prior 3.x era verified claims by lookup and arithmetic against consolidated GAAP series, which force-fit non-GAAP, segment, KPI, and qualitative claims to the wrong data; each added guardrail converted a wrong answer into no answer rather than fixing the root cause. We retired it in full and restarted on retrieval with verbatim sourcing. The error catalog from that era, including dated voids with pre-adjustment verdicts preserved, stays public: a clean-looking backfill is evidence of a contaminated instrument, and our mistakes are the certificate that this one ran blind.
Amendment, 2026-08-06: restated commitments, an announcement sentence and then the specific bar, are marked by a dedup agent with a written reason per merge, and one commitment counts once everywhere: scoring, counts, artifacts, the sealed record, and this site. Published n and accuracy moved when this landed; grades on restatements remain in the database, auditable, and merges are reversible. Declared here, never quietly retuned.
The record you can check
Every publish day the full score state is frozen and sealed with a SHA-256 digest over a canonical serialization, listed on the records page and archived externally the same day. Corrections never edit a grade in place: they enter a dated human-resolution lane that preserves the original verdict, and every database update to a grade is captured by an append-only audit trigger. Records sealed before the public launch are backfill and are labelled as such. If you believe a grade is wrong, the Disagree button on its claim page files it for review on the record.
Data challenges, named
Fiscal-period language is normalized to model-stated period integers, because exact-string labels fragment. Claims that lean on a prior quarter's guidance without restating it are graded only when the referent is recoverable, else abstained. Cost guidance inverts the deviation sign so favorability is consistent. Combined-company commitments are not graded against post-separation results. Where a company guides only on its own preferred basis, the grade says so.
Scores are opinions derived from public statements via this published methodology. Sealed digests: records.