· Court Transcript Platform
What our benchmark actually says
We ran our pipeline against the Supreme Court's own certified transcripts — and found our own benchmark had been scoring the better tier as the worse one. Here are the corrected numbers, now including the three-engine tier.
Update, 2026-08-12. The harness has moved on since this post: scoring now reconciles spelling variants to the certified orthography for every contender, excludes reporter's typography and verbatim policy from both sides, and includes Rev.com baselines — so every accuracy figure below is superseded. Current numbers live in Court transcription, by the numbers and the leaderboard. The figures below are kept as published: the story of finding our own benchmark's mistake is the point of this post, and that story is told in the numbers it was found in.
Most transcription vendors quote an accuracy percentage with no method behind it. No corpus, no ground truth, no way to check it.
Here is ours, including the parts that do not flatter us — starting with the part where the benchmark itself was wrong.
What we ran
Two U.S. Supreme Court oral arguments from April 20, 2026, Nos. 25-197 and 25-466. Both are public record, so anyone can obtain the same audio and official transcripts we scored against.
Ground truth is the official Heritage Reporting transcript of each argument: a real certified transcript, not crowd-sourced labels. Together they hold 24,577 words.
Every tier ran on that audio: one engine, two engines, and — new since the last version of this post — all three.
The mistake we found
Our reference transcripts are clean verbatim. Court reporters drop um and
uh; the official transcripts of these two arguments contain three filler words
between them.
Our engines are not all verbatim. We measured what each one returns on identical audio and found one of our two commercial engines returns almost no fillers at all — 2 across 24,577 words, against 248 from the other engine on the same recordings. Its vendor documents a flag to keep disfluencies. We set that flag on every request. It does not work.
That combination quietly broke the comparison. Our Good tier runs the engine that drops fillers. Our Advanced tier adds the engine that keeps them. Scored against a reference that also drops them, Advanced took an error for every filler it correctly heard — and Good never did.
So our own benchmark reported that the second engine bought nothing. It had been buying accuracy the whole time, and we were charging it for being faithful.
The fix is to report two error rates: one against the clean reference as-is, and one with fillers removed from both sides. Neither alone is honest. The first is what a court receives; the second is what the engines actually heard.
The results
| Hearing | Ground truth | Tier | Word accuracy | Filler-agnostic | Attribution |
|---|---|---|---|---|---|
| 25-197 | 11,309 words | Good | 94.8% | 94.8% | 98.3% |
| 25-197 | 11,309 words | Advanced | 94.8% | 95.8% | 97.9% |
| 25-197 | 11,309 words | Extreme | 94.8% | 95.7% | 97.8% |
| 25-466 | 13,268 words | Good | 91.9% | 91.9% | 98.6% |
| 25-466 | 13,268 words | Advanced | 92.0% | 92.8% | 98.2% |
| 25-466 | 13,268 words | Extreme | 91.8% | 92.6% | 98.1% |
Across all 24,577 words: Good reaches 93.2% and Advanced 93.3% on raw word accuracy — nearly identical. On filler-agnostic accuracy, Good stays at 93.2% and Advanced reaches 94.2%.
That gap is the real one. A full point of word accuracy, about a 14% reduction in errors, consistent across both hearings, invisible until we stopped penalizing the tier that transcribes what was actually said.
Attribution runs the other way. Good is slightly but consistently better — 98.5% against 98.1% across both hearings. The second engine buys word accuracy and costs a little speaker precision. That trade is real and we are not going to average it away.
The third engine did not make the transcript better
This is the run we most expected to flatter us, and it did not.
Extreme adds a third engine to Advanced's two. Across both hearings it lands at 93.2% raw and 94.0% filler-agnostic — level with Good on the first number and two tenths of a point behind Advanced on the second. Attribution slips again, to 98.0%. Both hearings move the same direction, so this is not one bad recording.
A third opinion is not a third of a vote toward the truth. Two engines already agree on the overwhelming majority of words; the third mostly arrives where they already agreed, and where it disagrees it is right roughly as often as it is wrong. On this corpus it costs a vendor pass and returns nothing on word accuracy.
We are publishing the tier anyway, because what it does buy is not accuracy. It is knowing which words to doubt.
Attribution is what legal work turns on: right words in the wrong mouth puts testimony on the wrong witness. At 98.5%, roughly 1 word in 66 lands on the wrong side of a speaker boundary, nearly all at turn transitions.
Disfluency retention
This is the metric meant to catch an engine that quietly cleans up a verbatim record. On 25-466, where the reference retains three fillers, the Good tier kept none of them and Advanced kept two of three.
We are reporting that honestly rather than dressing it up: three tokens is not a measurement. It is a signal pointing the same direction as the 2-versus-248 count, and the count is the number to trust. A corpus of clean-verbatim references structurally cannot measure verbatim fidelity well, which is why we now compare engines against each other on the same audio rather than relying on the reference alone.
The number that decides your review time
Every word carries a confidence score, worth something only if it is calibrated — words scored 0.9 really are right about 90% of the time. An overconfident score walks a reviewer straight past the errors.
Advanced tier, 24,546 scored words across both hearings:
| Confidence | Share of words | How often those words were right |
|---|---|---|
| 0.9 to 1.0 | 90.2% | 99.2% |
| 0.8 to 0.9 | 2.7% | 93.4% |
| 0.7 to 0.8 | 0.9% | 91.6% |
| below 0.5 | 3.1% | 35.2% |
Nine words in ten arrive above 0.9 and 99.2% of those are correct. Words the system doubts are wrong about 65% of the time. The score tracks reality in both directions, which is what lets you read the marked words in a 71-minute hearing instead of all 13,000.
One caveat we can see in our own data: Advanced puts 770 words below 0.5 where Good puts 60. Part of that band is real doubt, and part is an artifact — a filler only one engine returned scores as half-agreement even though both engines heard the audio correctly. We are fixing that upstream, and it will move this table.
How much you actually review
| Tier | Words flagged (both hearings) | Share of the transcript |
|---|---|---|
| Good | 64 | 0.26% |
| Advanced | 800 | 3.3% |
| Extreme | 823 | 3.3% |
Advanced flags about twelve times as much as Good, because it flags every word the two engines disagree on. Higher flag count, higher catch rate, more review time. The tier says so before you pick it.
At worst you are attending to about 1 word in 30, each anchored to the audio. Click the flag, the snippet plays, the transcript scrolls to the word, you confirm it or type the right one.
Three engines, and the pile that nearly buried them
Extreme sits beside Advanced in that table — 823 flags against 800. It very nearly did not.
Our flagging rule was: if any engine disagrees, a human looks. That rule is correct with two engines, where a disagreement is a genuine 1-1 tie that nothing in the run can settle. Applied to three engines it produced 1,755 flags on the same audio — more than double Advanced, for a transcript that measures slightly worse.
The second engine had also been dragging every score down by arithmetic. We scored each word by confidence times the share of engines that agreed, so a word two of three engines got right scored 0.67 and fell below the tier's own threshold. Adding an engine mechanically flagged more words. That is a property of the divisor, not of the audio.
So we changed both. A word now needs a majority against it, not a single dissenter, and a word a majority agreed on carries full weight rather than a fraction shrinking with the roster. Disagreement flags fell from 533 to 149: those 384 are words where the third engine settled a 2-1 dispute that used to cost a reporter a listen.
Both rules are identical to the old ones at one and two engines — two engines can never outvote each other, so a 1-1 tie still goes to a human, and every Good and Advanced number in this post is unchanged. We verified that the boring way, by re-running the whole benchmark and diffing: byte-identical.
That is what Extreme buys. Not a better transcript — the same review time as Advanced, aimed at different words. Of the errors the certified reference proves are errors, Extreme's flags catch 37.8% against Advanced's 36.8%.
One number in the table above will move. These Extreme counts were measured with the flag threshold at 0.70, inherited from before the vote existed. Having measured the curve, we have since fitted it to 0.55, which projects to 790 flags — below Advanced, at the same improved catch rate. We will republish the table when that run has gone through the harness end to end rather than our arithmetic.
For transcripts you bring from another service, we target under 15 minutes from upload to a filing-ready PDF. That is a target, not a measurement. We will publish the measured version when the pilot produces one.
The tiers
| Tier | Engines | Turnaround | Review burden |
|---|---|---|---|
| Good | One, the fastest | Fastest | Standard, engine confidence only |
| Advanced | Two, with consensus | Roughly double | Higher flags, higher catch rate |
| Extreme | Every engine, with a majority vote | Slowest | Comparable to Advanced, better aimed |
All three tiers ran on this corpus. Every export records which tier ran, which engines ran, how many flags were open, and whether the document was rendered verbatim or clean verbatim. That provenance travels with the document.
What these numbers do not prove
- Supreme Court audio is easy audio. Well-miked, one speaker at a time. Real hearings have crosstalk, bad room mics, and mumbling. Expect attribution to read lower there.
- The Good tier is not a verbatim tier. Its engine returns 2 fillers per 24,577 words. If you need a strictly verbatim record — and in most courts you do — that is the Advanced tier, and it is now also the more accurate one.
- Three filler tokens is not a disfluency measurement. The direction is clear; the sample is not.
- No competitor baselines ran. All three contenders are ours. Incumbent baselines on the same corpus are next.
- Two hearings is two hearings. This is a regression gate and a tier comparison, not an absolute accuracy claim.
- Three engines is not a verdict on three engines. Ours are three commercial vendors on clean audio, where they largely agree. A third engine may well earn its place on the messy recordings this corpus does not contain — that is the case where a majority vote has something to settle.
- A 1-point catch-rate difference is not a mandate. 37.8% against 36.8% is measured, small, and taken on two hearings.
- We never certify anything. We build the certification page. The attestation and signature belong to a licensed human.
Our previous version of this post reported that the second engine bought nothing on word error. That was our measurement error, not the engine's, and this is the correction. This version adds the third engine and reports that it bought nothing on word error either — that one we believe, and it is the vote, not the engine, that earns Extreme its place. When the next run moves these numbers, we will say so here too.
Get early access while the beta is free.