· Court Transcript Platform
One leaderboard: every mode, every engine alone, Rev.com, and a human on the flags
Seven ways to transcribe the same two Supreme Court arguments — our three modes, extreme mode with its flags resolved by a human, each engine by itself, and Rev.com — ranked on one table by errors against the certified record.
Our previous posts each measured one thing at a time — the modes against each other, Rev.com on our benchmark, what review buys. This post puts everything on one table: every mode we sell, every engine we run — by itself, so you can see what the ensemble adds — the extreme mode with a human resolving its flags, and the incumbent. Same audio, same certified ground truth, same blind scoring for every row.
The setup, briefly
Two full U.S. Supreme Court oral arguments — Nos. 25-197 (71 minutes) and 25-466 (74 minutes) — scored against the official Heritage Reporting transcripts: 23,544 scored certified reference words, the court's own transcript used as the answer key. Every contender is lined up against that answer key word by word, by the same automated grading, which doesn't know which transcript is whose. The Rev.com rows are transcripts we bought from Rev.com and graded the identical way.
One error column, deliberately. Scored errors reconciles spelling variants of the same spoken word to the certified orthography and removes reporter's typography and verbatim policy (fillers, stutters) from both sides, for every contender alike: it measures hearing, not house style. We do not publish a letter-for-letter "verbatim errors" column against this reference, because the reference itself is not verbatim — the official transcript drops the "uh"s and stutters — so such a count would rank contenders by their distance from the reporter's cleanup style, not by what they heard. (The strict count still lives in every measurement report, available on request.)
One row needs defining. "Extreme + review of flagged words": our system marks ("flags") every word it is unsure about — engines disagreeing, low confidence, a risky name — right in the draft, with the audio one click away. This row is the extreme mode's draft after a human checks only those marked words — about one word in twenty-five — and fixes the wrong ones, touching nothing else. No reading or listening through the whole document. It is the number from What review buys, and this table is where it lives in context. It assumes the reviewer gets every marked word right, so it is the ceiling for flag-guided review — and the floor for our certified tier, whose licensed human reads the whole record, not just the marks.
The leaderboard
Combined corpus, 23,544 scored reference words. Fewer errors is better:
| Rank | Contender | Scored errors | Accuracy |
|---|---|---|---|
| 1 | Extreme + review of flagged words | 328 | 98.6% |
| 2 | Advanced mode — 2-engine ensemble | 665 | 97.2% |
| 3 | ElevenLabs Scribe, alone | 668 | 97.2% |
| 4 | Extreme mode — 3-engine ensemble | 670 | 97.2% |
| 5 | Good mode — AssemblyAI, alone | 701 | 97.0% |
| 6 | Rev.com | 761 | 96.8% |
| 7 | Speechmatics, alone | 866 | 96.3% |
"Alone" means that engine as the only engine in our pipeline — same audio preparation, same case-name seeding, same speaker detection — so the gap between a single-engine row and an ensemble* row is the ensemble's contribution and nothing else.
* Ensemble: running two or three speech engines on the same audio and having them vote word by word — where they disagree, somebody is wrong, and that word gets flagged for a human.
Per hearing, because accuracy is a property of the audio as much as the engine and one combined number hides that:
| Contender | 25-197 scored errors | 25-466 scored errors |
|---|---|---|
| Extreme + review of flagged words | 129 | 199 |
| Advanced mode | 243 | 422 |
| ElevenLabs Scribe, alone | 252 | 416 |
| Extreme mode | 236 | 434 |
| Good mode (AssemblyAI, alone) | 278 | 423 |
| Rev.com | 309 | 452 |
| Speechmatics, alone | 330 | 536 |
Reading the rows
Extreme + review of flagged words — 328 errors. Resolving only the flags erases 51% of the extreme draft's errors, because the flags are aimed: cross-engine disagreement finds the errors a single engine's confidence never would. Versus Rev.com: 57% fewer errors (328 vs 761 on the same accounting), for roughly 3.6% of the words reviewed. This is also the row the certified tier builds on, not the certified tier itself.
Advanced mode — 665. The top unaided row — by three errors, which on this corpus is a tie with the two rows below it, and we say so rather than rank-inflate. Versus Rev.com: 13% fewer errors (665 vs 761), and on the case-caption words a record can least afford — parties, counsel, court — it errs at 1.5% to Rev.com's 2.0%.
ElevenLabs Scribe, alone — 668. The best single engine on this corpus, now visibly so: it is the most verbatim engine (it kept 124 fillers the reference drops), and the old scoring charged it for that fidelity. Scored on hearing alone, it ties the ensembles on words. What it cannot do alone is aim review — see the ensemble section below.
Extreme mode — 670. The word tie with advanced is expected — the third engine's vote changes little on clean audio. What it buys is better-aimed flags: its review catches 54% of remaining errors versus advanced's 49% on 25-466, and of the errors Rev.com made, its flags route 67% and 75% to a human on the two hearings. That is why the review row above is built on extreme.
Good mode (AssemblyAI, alone) — 701. Our floor and our fastest mode, now ahead of Rev.com on both hearings (278 vs 309 and 423 vs 452). One caveat travels with it: its engine silently drops the "um"s and false starts, which the scored column rightly doesn't charge — but if you need a strictly verbatim record, and in most courts you do, that is what the ensemble modes are for.
Rev.com — 761. Ahead of one of our seven rows (Speechmatics alone), behind the other six. Its most visible failures are names: its 25-466 export opens by calling the petitioner "Sweepich" — the case is Sripetch v. SEC — and misspells arguing counsel "Geiser" throughout. On caption words it errs at 2.0%, our two-engine mode at 1.5%. Of the errors Rev.com made on this corpus, our extreme mode's flags would have routed two thirds to three quarters to a human before anything shipped.
Speechmatics, alone — 866. Included for honesty: one of our three engines, by itself, loses to Rev.com. We run it anyway because ensembles want diverse voters, not three copies of the best one — its disagreements help aim the flags even where its own draft trails. Versus Rev.com: 14% more errors, the only row above that can say so.
What the ensemble is worth
The single-engine rows exist to answer one question: how much of our result is the engines, and how much is ours? The honest answer changed when the scoring got fair: on clean audio, the ensemble's word gain over the best rented engine is now a wash — 665 to Scribe's 668. The engines have converged on what clean audio can give.
What has not converged is knowing which words to doubt. A single engine flags nearly blind: the good mode's flags catch about 10% of its own errors, because with one voice there is no disagreement to detect. The ensemble's flags catch about half — and that difference is the entire distance between row 3 and row 1. Add the flags and a human pass over one word in twenty-five: 328 errors, a 51% cut below anything unaided on the table. The engines are rented; that conversion of review minutes into accuracy is the product.
The fine print, voluntarily
Same honesty box as every benchmark post. This corpus is clean, well-miked appellate audio — the ceiling condition, kinder to every row than a noisy municipal courtroom. The review row assumes a flawless reviewer; real ones err. Rev.com's rows are its standard automated product bought at retail, not its human tier — its human tier would score higher, cost more, and take longer; ours appears here only as the flags row for the same reason. The scored column removes spelling and verbatim policy from both sides for every contender, so nobody wins on house style; the letter-for-letter counts stay in the measurement reports. All numbers come from the same benchmark our CI runs on every release; the per-hearing measurement reports are available on request. Bring us audio that reorders the table.