· Court Transcript Platform
Court transcription, by the numbers
We benchmarked our pipeline against Rev.com on two U.S. Supreme Court arguments, scored against the official certified transcripts. Fewer word errors before review, less than half after it. Fewer name errors. Full measurement reports available on request.
Two U.S. Supreme Court oral arguments. 23,544 scored words of certified ground truth — the court's own official transcript, the answer key every contender is graded against. The same audio through our pipeline and through Rev.com, graded by the same automated scoring, which doesn't know which transcript is ours. The grading only counts hearing mistakes, never style: different spellings of the same spoken word ("7th"/"Seventh", "Rooker Feldman"/"Rooker-Feldman") are matched to the official spelling for every contender, Rev's included; punctuation the court reporter types (the bare "--" marking an interruption) is not counted as a word; and the "uh"s, cut-off syllables, and stuttered repeats ("I I don't") that the official transcript leaves out — but a faithful word-for-word engine writes down — are set aside on both sides. Nobody wins or loses on formatting policy; the scoreboard measures who heard the words right.
Here is what came out.
Word accuracy: the draft, and the draft after review
Word accuracy = the percentage of the certified transcript's words our transcript got right. Higher is better. Three columns, because our product is not a raw machine dump: the draft arrives with the words the system is unsure about already marked ("flagged"), and the middle column is that same draft after a human checks only the marked words — about 1 word in 25, each with its audio one click away — and fixes the wrong ones. That is the transcript that actually leaves the building. (It assumes the reviewer gets each marked word right — the measured ceiling of the flag system, not a promise about any particular reviewer.)
| Hearing | Ours, machine draft | Ours, after review of flagged words | Rev.com (as delivered) |
|---|---|---|---|
| No. 25-197 | 97.9% word accuracy | 98.8% | 97.2% |
| No. 25-466 | 96.6% | 98.3% | 96.4% |
Before review our best tier makes 24% fewer word errors on 25-197 and 7% fewer on 25-466 — close there, and we say so. After review it is not close: less than half the word errors Rev delivers (58% and 52% fewer), for a human pass over only about 3% of the words.
Fewer name errors
Names are what legal transcripts turn on. On the words from the case caption* — parties, counsel, the court — across hearing 25-466 (fewer is better):
| Errors on caption words | |
|---|---|
| Ours (ensemble) | 16 of 1,077 — 1.5% |
| Rev.com | 22 of 1,077 — 2.0% |
And on hearing 25-197 the ensemble makes 9 caption-word errors to Rev's 11.
Rev's export misspells the petitioner Sripetch as "Sweepich" in its opening line, and arguing counsel Geyser as "Geiser" throughout. Ours doesn't.
* The case caption is the heading of a court filing — who is suing whom, before which court, argued by which lawyers. Its words are the names a transcript can least afford to get wrong.
Who said it
Right words in the wrong mouth put testimony on the wrong witness. Scored with cpWER* on hearing 25-197:
| cpWER (lower is better) | |
|---|---|
| Ours | 0.054 |
| Rev.com | 0.065 |
That is roughly 17% fewer words affected by speaker mix-ups. On the second hearing Rev edges ours — 0.054 to our 0.055 — a number that does not flatter us, published anyway: a benchmark you can't lose is not a benchmark. (Numbers updated 2026-08-13, after our crosstalk re-attribution pass shipped — where two voices meet, a language model now re-sorts the words among the speakers our detection measured as present, and a human confirms every move; the measurement story is in Rev on our benchmark.)
* cpWER is the standard measure for "who said it": it counts a word as wrong not only when it is misheard, but also when it lands under the wrong speaker's name — and when a transcript splits one speaker in two or merges two into one, every word in the affected turns counts against it. 0.054 means about 5.4 words per 100 are touched by such a mistake. Lower is better.
The transcript tells you where to look
A number no vendor quotes: of the errors Rev.com provably made on this corpus, our review flags route about seven in ten to a human before certification (70% and 75% on the two hearings) — while asking them to look at only about 3% of the words. That routing is what turns the first table's draft column into its after-review column. The same uncertain name is one question, answered once, applied everywhere. Every certified transcript still gets a licensed human's final pass; the flags just aim it.
Check us
The corpus is public record: the hearings and their official certified transcripts are freely available, and Rev's exports were purchased like any customer would. Our full measurement reports are available on request, and the same benchmark runs on every release.
The long version — methodology, the metrics we invented nothing for, and the numbers that don't flatter us — is in Rev on our benchmark. We publish those too: a benchmark you can't lose is not a benchmark.