· Court Transcript Platform
How accurate is it? Here are the numbers.
Word accuracy, speaker attribution, and what human review buys — measured on certified court transcripts of public hearings. No adjectives, just the figures.
Every transcription product says "accurate." Almost none say a number, and fewer say how they got it. Here are ours: measured against certified official transcripts of two full U.S. Supreme Court arguments — 23,544 scored words of real courtroom ground truth, the court's own transcript used as the answer key — by the same automated scoring we run on every release. The scoring counts hearing mistakes, never style: different spellings of the same spoken word are matched to the official spelling for every contender, and the punctuation and "uh"s a court reporter leaves out are set aside on both sides.
The words
Word accuracy = the share of the official transcript's words we got right. Higher is better.
| Word accuracy | |
|---|---|
| Hearing 1 (71 minutes) | 97.9% |
| Hearing 2 (74 minutes) | 96.6% |
Two hearings, two numbers, not averaged: accuracy is a property of the audio as much as the engine, and one number for all audio is somebody's best day.
Who said it
The number that legal work actually turns on, and the one most services never publish — of the words we transcribed correctly, how many are also credited to the right speaker. Higher is better:
| Speaker attribution accuracy | |
|---|---|
| Hearing 1 | 98.3% |
| Hearing 2 | 98.7% |
Every speaking turn the system doubts — shaky audio, people talking over each other, or our speaker-detection passes disagreeing with each other — is queued for human confirmation before a transcript can certify. Doubt doesn't ship.
Head to head, same audio, same ground truth
We also bought transcripts of both hearings from a leading transcription service and graded them with the identical scoring, against the identical answer key. The scoring doesn't know which transcript is whose.
| Ours | Leading service | |
|---|---|---|
| Word accuracy (higher is better), hearing 1 | 97.9% | 97.2% |
| Word accuracy, hearing 2 | 96.6% | 96.4% |
| Errors on case names (fewer is better) | 25 | 33 |
That is 24% fewer word errors on the first hearing, ahead on both — and 24% fewer errors on names, the words a legal record can least afford to get wrong. Their export misspells the petitioner's name in its opening line and the arguing counsel's name throughout; ours doesn't.
One more number of that kind: of the errors that service made on this corpus, our review flags route about seven in ten to a human before anything certifies.
What review buys: 98.6%, reviewing one word in twenty-five
Our engines cross-check each other, and every word they disagree on, every word they are unsure of, and every high-consequence term — names, amounts, dates, citations, even where all engines agree — is marked in the draft ("flagged") with its audio snippet one click away. Reviewing means checking exactly those marked words, not reading the whole document; a recurring uncertain name is one question, answered once, applied everywhere.
Measured against the certified references: resolving the flags alone — about 1 word in 25, no listen-through — erases 51% of all remaining errors, taking the combined corpus to 98.6% word accuracy. The certified tier then adds what no automation replaces: a licensed human's final pass over the whole record, on top of that floor.
Confidence that means something
Every word carries a confidence score between 0 and 1 — the system's own estimate of how likely it heard that word right. "Calibrated" means the estimate is honest: words scored 0.9 really are right about nine times in ten, so the score can be trusted to aim a reviewer's attention:
| Confidence band | Share of words | Actually correct |
|---|---|---|
| 0.9 – 1.0 | 90.2% | 99.3% |
| below 0.5 | 3.2% | 36.1% |
Nine words in ten arrive above 0.9 and those are right 99 times in 100; the words the system doubts are wrong most of the time. That is what lets a reviewer read the marked words instead of all 13,000 — the score tracks reality in both directions.
The record stays verbatim
Court records keep the "um." General-purpose engines quietly clean speech — on identical audio we measured one engine writing down 2 filler words ("um", "uh") where another wrote 248. Our pipeline is built never to clean testimony: the "um"s, false starts, and repeats of natural speech are part of the record, silence that tempts an engine to invent words is structurally excluded, and no automated step may remove a spoken word.
The fine print, voluntarily
These figures come from clean, well-miked appellate audio — the ceiling condition, not a promise about a noisy municipal courtroom, where every vendor's numbers fall. Ground truth is the official certified transcript, not crowd labels. The corpus is public record — the hearings, their official transcripts, and our full measurement reports are available on request, and we run the same benchmark on every release. Bring us a hearing that breaks them.