What AI genuinely does well
On a single clear voice dictating at a moderate pace into a decent microphone, modern recognition is very good and very cheap. It is faster than any human, available immediately, and costs cents per minute. For a clinician who wants a draft to edit themselves, that is the right tool and a review stage would add cost for no benefit.
It is also good at scale in a way people are not. A practice processing hundreds of recordings a week gets a first pass on all of them in minutes, which is a genuine change in what is operationally possible.
Where it fails, and how the failures look
The failure modes are well understood and they are not random. Accented and rapid speech degrade accuracy. Two people talking over each other degrades it further. Ambient noise, speakerphones and cheap microphones compound with both. Specialist vocabulary - drug names, anatomical terms, laterality - fails at a lower rate but with higher consequence.
The characteristic that makes this a review problem rather than an accuracy problem is fluency. When older software failed, it produced obvious nonsense and you knew. When a model fails, it produces a phonetically similar drug name, a dose normalised toward a more usual value, an inferred laterality, or a dropped negation. Each reads correctly. Each is wrong in a way that survives a quick scan and reaches the chart.
- A similar-sounding drug substituted for the right one
- A dose quietly normalised toward a more common value
- Laterality inferred where the dictation did not state it
- A negation dropped, inverting the clinical meaning
The documentation risk AI introduced
Medicare imposes no prohibition on drafting documentation with software. What it requires is documentation supporting the billed service, and entries authenticated by whoever evaluated the patient. AI documentation is nonetheless disproportionately exposed, for three reasons: volume, because more notes are generated; fluency, because errors do not announce themselves; and timestamps, because these tools record exactly when a note was produced and signed.
That last point is the one practices underestimate. A system that records a note signed seconds after an encounter closed has created evidence that no meaningful review occurred. The only affirmative defence against a negligent-documentation claim involving AI-generated content is a contemporaneous, tamper-evident log showing the clinician reviewed the specific contested content before signing.
So what should a practice buy
Not "AI" or "human" as a blanket choice - the right answer differs per document. Unreviewed output for material you will read and correct yourself. Reviewed output for anything entering a chart or going to a third party. And, for the second category, a record that the review happened.
ScribeForms is built around that distinction rather than against it. Verification is chosen per job at a stated rate; extraction returns per-field confidence and the supporting quote so review attention goes where the model was least certain; and every approval is recorded over a digest of exactly the content the signer saw, with the elapsed time between transcript and signature. The AI does the draft. The record proves what happened next.
The measurement problem behind every accuracy claim
Vendors publish a single figure without a methodology, a corpus, or an error taxonomy. The number is unverifiable by construction: it is an average over audio somebody else selected, scored by a method they did not publish, and it tells a buyer nothing about the field they actually care about.
Consider what 99% means on a 1,500-word consultation. Roughly fifteen errors. Whether that is excellent or dangerous depends entirely on where they fall, and an average cannot tell you. A per-field confidence score can: it says which values the system was least sure about, so a reviewer looks there first rather than reading the whole document with a uniform attention nobody can sustain.
This is why we publish confidence per field and the transcript quote behind each extracted value, and publish no headline accuracy figure at all. The checkable claim is worth more than the bigger one, and the buyers in medical and legal work are precisely the ones who check.
Where the human stage still earns its cost
On clean single-speaker audio with ordinary vocabulary, review adds little. On accented or rapid dictation, on two-party consultations with crosstalk, on dense specialist terminology, and on anything where a figure or a laterality matters, it recovers errors an automated pass will not. The distinction is predictable enough that a practice can decide per document rather than per contract.
That is the argument for making verification a per-job choice rather than a property of the account. A clinic sending both a clinical dictation and a staff meeting through the same service should not pay for review twice, and should not have to go without it once.
Sources
- 42 U.S.C. § 1395l(e) (documentation supporting the billed service)
- 42 C.F.R. § 482.24(c)(1) (authentication by the clinician responsible)
- 45 C.F.R. § 164.312(b) (audit controls)
Verified 19 September 2026.
The regulatory information on this page is general background compiled from public primary sources, not legal or compliance advice. Requirements change and vary by jurisdiction and by court. Verify current rules with the relevant authority or your own counsel before relying on them.