· Court Transcript Platform
Six hearings, by the numbers
Our benchmark corpus tripled: six U.S. Supreme Court arguments, 8.4 hours of audio, 89,328 certified words. Every error graded for whether it changes the record, and every one counted as caught or missed by our review flags. The numbers, including the ones that don't flatter us.
When we last published numbers, the benchmark was two Supreme Court arguments. It is now six — Nos. 24-1068, 25-112, 25-197, 25-406, 25-429 and 25-466, argued in April 2026 — 8.4 hours of audio and 89,328 scored words of certified ground truth, the Court's own official transcripts used as the answer key. Same blind grading as before: spelling variants of the same spoken word are reconciled to the official spelling for every contender, the reporter's typography is not counted as a word, and the fillers and stutters a clean-verbatim transcript leaves out are set aside. The scoreboard measures who heard the words right, not house style.
Two things are new in how we count. First, every error is graded for whether it changes what the record says — "serious" — or leaves the meaning intact — "minor": a dropped article, a figure written in digits instead of words, a name split in two. Second, every error is marked caught or missed: caught if one of our review flags points a human at it, missed if it would sail through flag-guided review untouched. A benchmark that only reports a word error rate hides both of those, and they are the two numbers a court actually cares about.
All figures below are our advanced mode — the two-engine ensemble that is our default — unless a column says otherwise.
Word accuracy, hearing by hearing
Word accuracy = the share of the certified transcript's words the draft got right. The middle column is the same draft after a human checks only the flagged words — about 1 word in 30 — and fixes the wrong ones; it assumes the reviewer gets each flagged word right, so it is the measured ceiling of the flag system, not a promise about any particular reviewer — computed by removing the share of errors the flags reach from the draft's error rate.
| Hearing | Audio | Scored words | Advanced, machine draft | Advanced, after review of flags | Good mode, draft |
|---|---|---|---|---|---|
| No. 25-197 | 62 min | 10,999 | 97.8% | 98.7% | 97.5% |
| No. 25-429 | 90 min | 15,078 | 97.1% | 98.1% | 96.0% |
| No. 25-406 | 85 min | 15,536 | 97.0% | 98.5% | 95.8% |
| No. 25-466 | 71 min | 12,542 | 96.8% | 98.4% | 96.7% |
| No. 24-1068 | 74 min | 12,927 | 96.5% | 98.5% | 94.6% |
| No. 25-112 | 120 min | 22,246 | 96.4% | 97.8% | 95.8% |
| All six | 8.4 h | 89,328 | 96.9% | 98.6% | 96.0% |
Across the corpus the draft charges 2,784 word errors in 89,328 words — 3.1%. Another 1,973 differences were forgiven as style before counting: "Seventh" for "7th", a hyphen the reporter typed, an "uh" the engine faithfully wrote down. They are in the reports; they are not errors.
The spread is the honest part. The two hearings we published last month are still the easy end; the four new ones are longer, have more advocates, and 25-112 — two hours of a Fourth Amendment argument dense with case names — is the hardest audio we have measured.
Which errors matter
Those word errors group into 1,949 issues (an issue is one run of disagreement with the record — a dropped clause is one issue, several words). Of those:
| Hearing | Issues | Serious — changes the record | Minor — meaning survives | Meaningful error rate* |
|---|---|---|---|---|
| No. 25-197 | 156 | 41 | 115 | 0.76% |
| No. 25-429 | 304 | 61 | 243 | 0.66% |
| No. 25-406 | 350 | 99 | 251 | 0.91% |
| No. 25-466 | 302 | 72 | 230 | 0.89% |
| No. 24-1068 | 288 | 60 | 228 | 0.91% |
| No. 25-112 | 549 | 165 | 384 | 1.16% |
| All six | 1,949 | 498 | 1,451 | 0.9% |
One error in four changes what the record says. The rest are the kind a reader corrects without noticing — real errors, counted in the 3.1%, but not the ones that put the wrong word in a witness's mouth.
* Word error rate counting only the errors graded serious. Roughly: about nine words per thousand in the draft would change the record if left alone.
Caught or missed
This is the table no vendor publishes. For each hearing: how many words we asked a human to look at, and of the serious and minor issues, how many a flag reached.
| Hearing | Flagged words per 1,000 | Serious issues caught | Minor issues caught | Serious issues that get past review |
|---|---|---|---|---|
| No. 25-197 | 31 | 28 of 41 — 68% | 46 of 115 | 13 |
| No. 25-429 | 35 | 43 of 61 — 70% | 128 of 243 | 18 |
| No. 25-406 | 35 | 44 of 99 — 44% | 144 of 251 | 55 |
| No. 25-466 | 36 | 53 of 72 — 74% | 120 of 230 | 19 |
| No. 24-1068 | 37 | 28 of 60 — 47% | 145 of 228 | 32 |
| No. 25-112 | 34 | 80 of 165 — 48% | 166 of 384 | 85 |
| All six | 34 | 276 of 498 — 55% | 749 of 1,451 — 52% | 222 |
Read it plainly: the flags route a little over half of the errors that matter to a human, for a review pass over 3.4% of the words. The other 222 serious issues across 8.4 hours — about one every 400 words — ship in the draft and are only removed by a full read-through, which every certified transcript still gets. We would rather you knew the size of that residue than took "98.5% after review" at face value.
The residue is not evenly spread. On 25-466, 25-197 and 25-429 the flags reach seven in ten serious issues; on 25-406, 24-1068 and 25-112 fewer than half. What those three share is that both engines agreed on the wrong word more often — a consensus error, invisible to any flag built on disagreement or confidence. That is the ceiling of flagging as a technique, and it is where the remaining work is.
The good mode, for comparison, flags 15–24 words per thousand and reaches only 5–19% of its own errors: a single engine's confidence is a poor guide to where it is wrong. The ensemble's disagreement is the signal.
Names and caption terms
On the words from each case caption — parties, counsel, the court — across all six hearings:
| Caption-term errors | …of which change the meaning | |
|---|---|---|
| Advanced mode, draft | 112 of 6,172 — 1.8% | 36 — 0.6% |
Most of the 112 are a name split or hyphenated differently from the reporter's spelling; 36 are a wrong name. Speaker identification — putting the right name on each voice, from the caption and the argument's own "Mr. Chief Justice, and may it please the Court" — named 50 of 54 speaker clusters correctly across the corpus.
Who said it
Right words in the wrong mouth put testimony on the wrong witness. cpWER* per hearing, lower is better:
| Hearing | cpWER | Turn attribution accuracy |
|---|---|---|
| No. 25-429 | 0.053 | 98.6% |
| No. 25-197 | 0.055 | 98.1% |
| No. 25-466 | 0.058 | 98.5% |
| No. 25-406 | 0.063 | 98.1% |
| No. 24-1068 | 0.067 | 98.3% |
| No. 25-112 | 0.072 | 97.8% |
* cpWER counts a word as wrong when it is misheard or lands under the wrong speaker; a speaker split in two or two merged into one costs every word in the affected turns. 0.053 means about 5.3 words per 100 are touched by such a mistake.
Against Rev.com, where we have it
We hold Rev.com transcripts for the two original hearings, bought like any customer would. On those two, graded identically:
| Ours, advanced draft | Rev.com, as delivered | |
|---|---|---|
| Word errors, No. 25-197 | 240 (97.8%) | 326 (97.0%) |
| Word errors, No. 25-466 | 396 (96.8%) | 445 (96.4%) |
| Serious issues, No. 25-197 | 41 | 91 |
| Serious issues, No. 25-466 | 72 | 141 |
| Meaningful error rate | 0.76% / 0.89% | 1.34% / 1.55% |
| Caption-term errors (both) | 27 | 34 |
Fewer errors overall, and about half as many of the errors that change the record — before any human looks at our draft. Speaker attribution remains close: Rev edges us on 25-466 (cpWER 0.054 to our 0.058), we edge Rev on 25-197 (0.055 to 0.065). We have not bought Rev transcripts for the four new hearings; when we do, they go on this table whichever way they fall.
What changed under the hood, and how we kept ourselves honest
Since the last post a small learned ranker joined the review flags: it scores the words no rule flagged and spends a bounded extra budget — about three more flagged words per thousand — on the highest scorers. It is trained on this public corpus, never on customer transcripts. Which creates an obvious temptation: quote its recall on the hearings it learned from. We don't. One hearing, No. 25-466, is held out of training entirely, and it is the only hearing whose numbers we will quote with the ranker switched on: serious-and-minor error recall there goes from 49.9% to 54.2% at 3.3 more flags per thousand. The tables above are the rule flags alone, so every hearing is on an equal footing.
We also tried the obvious next steps — a gradient-boosted model in place of the logistic regression, sixteen context features, the dissenting engines' own proposals, and a language model's surprise at each word — and none moved recall beyond the run-to-run noise of the pipeline (about one percentage point). We publish that too. Per-word ranking has found its ceiling on this corpus; the consensus errors need a different tool.
Check us
The corpus is public record: the six arguments are on the Supreme Court's own site (supremecourt.gov, Argument Audio, October Term 2025) and the official transcripts beside them. Our per-hearing measurement reports — every charged error with its context, its grade, and whether a flag reached it — are available on request, and the same benchmark runs on every release.
A benchmark you can't lose is not a benchmark.