Stage one: the recording
Everything downstream is bounded by this. A clear single voice close to a decent microphone produces a transcript that needs little correction; a speakerphone in a room with an air conditioner produces one that needs a lot. The difference in output quality between a good and a poor recording is larger than the difference between any two transcription providers.
Practical consequences: dictate in a quiet room, hold the device consistently, and state names and figures slowly and once rather than quickly and twice. If the recording captures a two-party consultation rather than a dictation, say who is who at the start - a provider that can identify speakers still cannot know their names unless the recording contains them.
Stage two: upload and the disclosure it creates
The audio leaves the practice, and at that moment protected health information has been disclosed to a third party. Under 45 CFR 164.308(b)(1) that makes the provider a business associate, and the agreement must be in place before the disclosure rather than after it. This is the stage practices handle most casually and the one with the clearest legal requirement.
Operationally, upload is also where file limits bite. Transcription APIs have size caps, and a long recording has to be split and stitched back together. A provider that does this properly preserves the timings across the joins; one that does not produces a transcript where clicking a line seeks to the wrong moment in the audio.
Stage three: the automated draft
Speech recognition produces text with word-level timings, and where the provider supports it, speaker labels. This takes minutes rather than hours and costs cents rather than dollars. It is also where the errors enter, and they enter in a specific shape: fluent, plausible, and concentrated in exactly the content that matters - drug names, doses, laterality, negations.
A provider worth using will tell you which parts of the draft it was least confident about. One that returns a flat block of text with a single accuracy claim has given a reviewer no way to prioritise, which means the review is either exhaustive and expensive or cursory and ineffective.
Stage four: review, if you bought it
A reviewer listens against the draft and corrects it. On clean audio with ordinary vocabulary there is not much to do. On accented speech, crosstalk or dense terminology this is where the value is, and it is the stage that costs money because it costs somebody's time.
The question to ask is not whether a provider has a review process but what evidence of it you receive. "We have QA" is unverifiable. A record naming who reviewed the document, over a digest of exactly what they saw, at a stated time, is not.
Stage five: formatting, delivery and sign-off
The transcript becomes a document: sections in the order the document type calls for, speakers attributed, timings available if the output needs captions. Delivery is a file, an email, or - better - a webhook or API call that puts the document into your own system without anyone retyping it.
Then a clinician approves it, and that approval is the step with the most legal weight. 42 CFR 482.24(c)(1) requires entries to be authenticated by the person responsible for the service. CMS audit posture now treats a note signed seconds after an encounter closed as evidence that no review occurred - so the timing of this step is itself part of the record.
What this looks like on ScribeForms
Record or upload in the browser; choose per job whether a person reviews it, with the rate difference shown before you submit rather than discovered on an invoice. The draft comes back with per-field confidence and the transcript quote behind each extracted value, so review attention goes where the model was least certain.
If the document you actually need is a form rather than a transcript, upload the blank form and dictate the answers - we write them into that same document in its own layout. And when you approve, we record who you are, a SHA-256 digest of exactly the content you saw, the version, the time, and the interval between the transcript becoming available and your signature. That certificate exports, and a recipient can verify the digest without us.
Sources
- 45 C.F.R. § 164.308(b)(1) (business associate agreement required before disclosure)
- 42 C.F.R. § 482.24(c)(1) (authentication by the responsible clinician)
- 45 C.F.R. § 164.312(b) (audit controls)
Verified 19 September 2026.
The regulatory information on this page is general background compiled from public primary sources, not legal or compliance advice. Requirements change and vary by jurisdiction and by court. Verify current rules with the relevant authority or your own counsel before relying on them.