§Court Transcript Platform
← All posts

· Court Transcript Platform

Six hearings, by the numbers

Our benchmark corpus tripled: six U.S. Supreme Court arguments, 8.4 hours of audio, 89,328 certified words. Every error graded for whether it changes the record, and every one counted as caught or missed by our review flags. The numbers, including the ones that don't flatter us.

When we last published numbers, the benchmark was two Supreme Court arguments. It is now six — Nos. 24-1068, 25-112, 25-197, 25-406, 25-429 and 25-466, argued in April 2026 — 8.4 hours of audio and 89,328 scored words of certified ground truth, the Court's own official transcripts used as the answer key. Same blind grading as before: spelling variants of the same spoken word are reconciled to the official spelling for every contender, the reporter's typography is not counted as a word, and the fillers and stutters a clean-verbatim transcript leaves out are set aside. The scoreboard measures who heard the words right, not house style.

Two things are new in how we count. First, every error is graded for whether it changes what the record says — "serious" — or leaves the meaning intact — "minor": a dropped article, a figure written in digits instead of words, a name split in two. Second, every error is marked caught or missed: caught if one of our review flags points a human at it, missed if it would sail through flag-guided review untouched. A benchmark that only reports a word error rate hides both of those, and they are the two numbers a court actually cares about.

All figures below are our advanced mode — the two-engine ensemble that is our default — unless a column says otherwise.

Word accuracy, hearing by hearing

Word accuracy = the share of the certified transcript's words the draft got right. The middle column is the same draft after a human checks only the flagged words — about 1 word in 30 — and fixes the wrong ones; it assumes the reviewer gets each flagged word right, so it is the measured ceiling of the flag system, not a promise about any particular reviewer — computed by removing the share of errors the flags reach from the draft's error rate.

HearingAudioScored wordsAdvanced, machine draftAdvanced, after review of flagsGood mode, draft
No. 25-19762 min10,99997.8%98.7%97.5%
No. 25-42990 min15,07897.1%98.1%96.0%
No. 25-40685 min15,53697.0%98.5%95.8%
No. 25-46671 min12,54296.8%98.4%96.7%
No. 24-106874 min12,92796.5%98.5%94.6%
No. 25-112120 min22,24696.4%97.8%95.8%
All six8.4 h89,32896.9%98.6%96.0%

Across the corpus the draft charges 2,784 word errors in 89,328 words — 3.1%. Another 1,973 differences were forgiven as style before counting: "Seventh" for "7th", a hyphen the reporter typed, an "uh" the engine faithfully wrote down. They are in the reports; they are not errors.

The spread is the honest part. The two hearings we published last month are still the easy end; the four new ones are longer, have more advocates, and 25-112 — two hours of a Fourth Amendment argument dense with case names — is the hardest audio we have measured.

Which errors matter

Those word errors group into 1,949 issues (an issue is one run of disagreement with the record — a dropped clause is one issue, several words). Of those:

HearingIssuesSerious — changes the recordMinor — meaning survivesMeaningful error rate*
No. 25-197156411150.76%
No. 25-429304612430.66%
No. 25-406350992510.91%
No. 25-466302722300.89%
No. 24-1068288602280.91%
No. 25-1125491653841.16%
All six1,9494981,4510.9%

One error in four changes what the record says. The rest are the kind a reader corrects without noticing — real errors, counted in the 3.1%, but not the ones that put the wrong word in a witness's mouth.

* Word error rate counting only the errors graded serious. Roughly: about nine words per thousand in the draft would change the record if left alone.

Caught or missed

This is the table no vendor publishes. For each hearing: how many words we asked a human to look at, and of the serious and minor issues, how many a flag reached.

HearingFlagged words per 1,000Serious issues caughtMinor issues caughtSerious issues that get past review
No. 25-1973128 of 41 — 68%46 of 11513
No. 25-4293543 of 61 — 70%128 of 24318
No. 25-4063544 of 99 — 44%144 of 25155
No. 25-4663653 of 72 — 74%120 of 23019
No. 24-10683728 of 60 — 47%145 of 22832
No. 25-1123480 of 165 — 48%166 of 38485
All six34276 of 498 — 55%749 of 1,451 — 52%222

Read it plainly: the flags route a little over half of the errors that matter to a human, for a review pass over 3.4% of the words. The other 222 serious issues across 8.4 hours — about one every 400 words — ship in the draft and are only removed by a full read-through, which every certified transcript still gets. We would rather you knew the size of that residue than took "98.5% after review" at face value.

The residue is not evenly spread. On 25-466, 25-197 and 25-429 the flags reach seven in ten serious issues; on 25-406, 24-1068 and 25-112 fewer than half. What those three share is that both engines agreed on the wrong word more often — a consensus error, invisible to any flag built on disagreement or confidence. That is the ceiling of flagging as a technique, and it is where the remaining work is.

The good mode, for comparison, flags 15–24 words per thousand and reaches only 5–19% of its own errors: a single engine's confidence is a poor guide to where it is wrong. The ensemble's disagreement is the signal.

Names and caption terms

On the words from each case caption — parties, counsel, the court — across all six hearings:

Caption-term errors…of which change the meaning
Advanced mode, draft112 of 6,172 — 1.8%36 — 0.6%

Most of the 112 are a name split or hyphenated differently from the reporter's spelling; 36 are a wrong name. Speaker identification — putting the right name on each voice, from the caption and the argument's own "Mr. Chief Justice, and may it please the Court" — named 50 of 54 speaker clusters correctly across the corpus.

Who said it

Right words in the wrong mouth put testimony on the wrong witness. cpWER* per hearing, lower is better:

HearingcpWERTurn attribution accuracy
No. 25-4290.05398.6%
No. 25-1970.05598.1%
No. 25-4660.05898.5%
No. 25-4060.06398.1%
No. 24-10680.06798.3%
No. 25-1120.07297.8%

* cpWER counts a word as wrong when it is misheard or lands under the wrong speaker; a speaker split in two or two merged into one costs every word in the affected turns. 0.053 means about 5.3 words per 100 are touched by such a mistake.

Against Rev.com, where we have it

We hold Rev.com transcripts for the two original hearings, bought like any customer would. On those two, graded identically:

Ours, advanced draftRev.com, as delivered
Word errors, No. 25-197240 (97.8%)326 (97.0%)
Word errors, No. 25-466396 (96.8%)445 (96.4%)
Serious issues, No. 25-1974191
Serious issues, No. 25-46672141
Meaningful error rate0.76% / 0.89%1.34% / 1.55%
Caption-term errors (both)2734

Fewer errors overall, and about half as many of the errors that change the record — before any human looks at our draft. Speaker attribution remains close: Rev edges us on 25-466 (cpWER 0.054 to our 0.058), we edge Rev on 25-197 (0.055 to 0.065). We have not bought Rev transcripts for the four new hearings; when we do, they go on this table whichever way they fall.

What changed under the hood, and how we kept ourselves honest

Since the last post a small learned ranker joined the review flags: it scores the words no rule flagged and spends a bounded extra budget — about three more flagged words per thousand — on the highest scorers. It is trained on this public corpus, never on customer transcripts. Which creates an obvious temptation: quote its recall on the hearings it learned from. We don't. One hearing, No. 25-466, is held out of training entirely, and it is the only hearing whose numbers we will quote with the ranker switched on: serious-and-minor error recall there goes from 49.9% to 54.2% at 3.3 more flags per thousand. The tables above are the rule flags alone, so every hearing is on an equal footing.

We also tried the obvious next steps — a gradient-boosted model in place of the logistic regression, sixteen context features, the dissenting engines' own proposals, and a language model's surprise at each word — and none moved recall beyond the run-to-run noise of the pipeline (about one percentage point). We publish that too. Per-word ranking has found its ceiling on this corpus; the consensus errors need a different tool.

Check us

The corpus is public record: the six arguments are on the Supreme Court's own site (supremecourt.gov, Argument Audio, October Term 2025) and the official transcripts beside them. Our per-hearing measurement reports — every charged error with its context, its grade, and whether a flag reached it — are available on request, and the same benchmark runs on every release.

A benchmark you can't lose is not a benchmark.