· Court Transcript Platform
What review buys, measured
If a human resolved every word our system flags — and nothing else — how accurate is the transcript? We computed it against certified references: 98.6% combined, reviewing one word in twenty-five. Here is that number, its cost, and what the residual actually is.
Every transcription service with a human tier quotes the same number: 99%. None of them tell you what the machine-directed part of that review is worth — how much accuracy the flags* themselves buy, before anyone reads the rest of the document.
* A flag is a mark the system puts on a word it is unsure about — the engines disagreed, the confidence was low, or the word is a risky name or number. Reviewing means checking exactly the marked words, each with its audio one click away — not reading or listening through the document.
We measured exactly that, because our benchmark can: take the machine's draft, assume a human checks every flagged word and fixes it correctly — and touches nothing that isn't flagged — and grade what remains against the court's own certified transcript. What survives is precisely the set of errors our system never surfaced.
The numbers
Two U.S. Supreme Court arguments, 23,544 scored certified reference words, our extreme tier. Accuracy: higher is better; errors: fewer is better:
| Hearing | Draft accuracy | After review of flagged words | Errors erased | Words reviewed |
|---|---|---|---|---|
| 71-minute argument | 97.9% | 98.8% | 236 → 129 | 2.9% |
| 74-minute argument | 96.5% | 98.4% | 434 → 199 | 4.1% |
| Combined | 97.2% | 98.6% | 670 → 328 (−51%) | ~3.6% |
Reviewing one word in twenty-five — click a flag, hear the snippet, pick the right reading — erases 51% of all transcription errors. No listen-through required for that gain; the flags aim every minute of it.
The ensemble earns its keep here, visibly: the single-engine tier's flags erase only ~10% of errors on the same audio, because with one engine there is no disagreement to detect. Cross-checking engines is what turns review time into accuracy.
What the residual 1.4% is
The errors that survive are the interesting ones: words every engine got wrong the same way, so no disagreement or confidence signal exists to flag them. Mostly hard names and terms the engines share blind spots on. This number is now printed by every benchmark run we do, because it is the cleanest possible progress bar for the work aimed at it — case-vocabulary biasing from uploaded filings, and the correction flywheel that makes every reviewer fix teach the next hearing.
What this number is not
Honesty box, as always:
- It assumes the reviewer resolves every flag correctly — an idealization. In practice most flag resolutions are one click between offered readings, but humans err.
- It is the flags-only floor, not our certified tier. A certified transcript additionally gets a licensed human's pass over the whole document, which catches unflagged errors too — real certified accuracy sits above this number. We publish the floor because it is the part we can measure without grading our own homework.
- Same corpus caveat as ever: clean, well-miked appellate audio. This is the ceiling condition; your county's motion calendar is harder, for every vendor.
The full measurement reports behind every figure are available on request, and this post-review number is now computed by every run of our benchmark — it is how we will track the residual shrinking.