Why accuracy percentages mislead
A vendor quoting "99% accuracy" is describing word error rate across some unnamed corpus, usually clean single-speaker audio recorded in good conditions. Your files are not that corpus.
More importantly, word error rate treats all words as equal. In a clinical note, a misheard conjunction and a misheard drug dosage both count as one error. In a deposition, a filler word and a misattributed answer count the same. The metric is indifferent to exactly the distinction that determines whether a transcript is usable.
A transcript that is 98% accurate with the 2% falling on articles and filler words is fine. One that is 99% accurate with the 1% falling on names, numbers, and speaker labels can be worse than useless, because it reads as authoritative while being wrong where it counts.
What AI transcription reliably gets wrong
These failure modes are consistent enough to plan around, which is the useful thing about them. They are not random degradation; they cluster in predictable places, and every one of those places happens to be where professional transcripts carry their risk.
- Proper nouns and unfamiliar terminology — a plausible-sounding substitute is produced rather than a flag raised.
- Overlapping speech and speaker changes — the point where multi-party audio degrades fastest.
- Numbers, dosages, and figures, which have no linguistic context to constrain the guess.
- Accented and non-native speech, and code-switching mid-sentence.
- Narrowband telephone audio, common in recorded statements and claims work.
- Homophones that are context-dependent rather than grammatically determined.
The silence problem
The structural issue with automated-only transcription is not the error rate. It is that the output carries no signal about its own reliability.
A human transcriptionist unsure of a word marks it inaudible or flags it. A model produces its best guess with no indication that a guess was made. The transcript reads uniformly confident, so a reader has no way to know which passages to verify against the audio — which means either verifying all of it, defeating the point, or trusting all of it, which is the actual risk.
This is why review is a stage rather than a feature. It is the only point at which uncertainty gets surfaced.
Where the two approaches are genuinely equivalent
Arguing that human review is always worth buying would be a sales position rather than an honest one. There is a real class of work where automated transcription is the correct choice, and pretending otherwise makes the rest of the argument less credible.
Single-speaker audio recorded in good conditions, on general subject matter, with no proper nouns that matter and no consequence to a quiet error — that is where models perform closest to their published benchmarks. Searching your own meeting recording for the part you half-remember, producing a rough draft you will rewrite anyway, or getting the gist of a long recording before deciding whether to transcribe it properly: buy the cheap version.
The distinction is not really about the audio. It is about whether anyone will rely on the transcript without re-listening. A transcript you will check against the source as you use it does not need to have been checked for you. A transcript that will be filed, charted, coded, or cited does.
Hybrid workflows, and the trap in them
A common middle path is to run everything through automated transcription and send only the important files for review. This works, with one qualification that tends to get discovered late.
It requires knowing which files are important before reading them, and that is frequently the thing you do not know. The recorded statement that turns out to contain the decisive qualification, the consultation where an unexpected finding was dictated, the deposition passage that becomes central months later — these are not identifiable in advance from file metadata.
Where the set of consequential files is knowable up front, triage is sensible and saves real money. Where it is not, triage amounts to deciding at random which of your records are allowed to contain unflagged errors.
A decision rule that holds up
Ask what a quiet error costs. If a wrong dosage, a misattributed answer, or a transposed figure would cost more than the price difference between reviewed and unreviewed transcription, buy the review. If the transcript is a convenience — searching your own meeting recording, drafting notes you will rewrite anyway — automated output is genuinely adequate.
Most professional work falls on the first side, which is why regulated fields converged on human review rather than adopting automated-only transcription when it became cheap. The cost of the review is small and predictable; the cost of the error is neither.