Why an AI-Generated Note Fails an Audit

5 min read

Why do AI-generated clinical notes fail audits?

Not because AI produced them. Medicare imposes no prohibition on drafting documentation with software, and an AI-assisted note that a clinician genuinely reviewed and signed is a valid record. Notes fail when the record cannot show that review happened — when the documentation does not support the level of service billed, when it is signed but not authenticated in a way that survives scrutiny, or when the only evidence of review is the clinician saying afterwards that they reviewed it. The exposure is evidentiary rather than technological, which is why it is fixable.

What the rules actually require

Two obligations do the work here, and neither mentions artificial intelligence. The first is that services billed to Medicare must be documented. 42 U.S.C. 1395l(e) conditions payment on the provider furnishing information necessary to determine the amounts due, and the Medicare Program Integrity Manual instructs reviewers to deny claims where documentation does not support the service billed. A note that reads well but does not substantiate the code is a denial regardless of who or what drafted it.

The second is authentication. 42 C.F.R. 482.24(c)(1) requires that all patient medical record entries be legible, complete, dated, timed, and authenticated by the person responsible for providing or evaluating the service. That is the provision an AI-drafted note interacts with, and the operative word is authenticated. An entry attributed to a clinician who did not actually evaluate its contents is not authenticated in the sense the regulation means, whatever the signature block says.

Put together, the requirement is not that a human wrote the note. It is that a human took responsibility for its contents, and that the record reflects this. Software that drafts is acting as a tool; software whose output enters the chart unexamined has effectively become the author, and there is no provision under which that is acceptable.

Why AI documentation is disproportionately exposed

The regulatory text is technology-neutral, but the practical risk is not evenly distributed, for three reasons specific to how these tools are used.

The first is volume. A tool that turns a consultation into a finished note in under a minute produces far more documentation per clinician than dictation did. Where review is perfunctory, the same defect is replicated across every encounter rather than appearing occasionally, which is exactly the pattern statistical review is designed to surface.

The second is plausibility. A transcription error produces something obviously wrong. A language model produces something fluent, internally consistent, and wrong — a laterality flipped, a dose plausible but not the one stated, a symptom the model expected in that clinical picture but which the patient never reported. Errors that read correctly are the ones that survive a fast review, and they are also the ones that matter most.

The third is that these tools generate their own timestamps. An encounter that closes and produces a signed note seconds later leaves a record of that interval in the system, and the interval is not consistent with anyone having read the content. Contemporaneous metadata is ordinarily a provider protection; here it can document the absence of the very review the signature asserts.

The defence is a record, not a policy

The common response is a written policy stating that clinicians must review AI-generated output before signing. That is worth having and it is not evidence. A policy establishes what was supposed to happen; an audit or a negligence claim asks what did happen in the specific encounter under examination, and a policy document says nothing about any particular note.

What answers that question is a contemporaneous, system-generated record showing that a named person reviewed specific content at a specific time, made or declined to make changes, and then signed. The distinction is the same one that makes a delivery receipt more useful than a shipping policy. One describes intent; the other records an event.

Three properties determine whether such a record is worth anything under challenge. It must be contemporaneous, created as part of the review rather than reconstructed afterwards from memory or from a spreadsheet. It must be tamper-evident, so that content cannot be altered after signing without the alteration being detectable — otherwise the record proves only that someone signed something, not what they signed. And it must be exportable, because a record that cannot leave the system that created it is not something you can attach to an audit response.

  • Contemporaneous — generated during review, not assembled later.
  • Tamper-evident — a cryptographic hash of the reviewed content, so post-signature changes are detectable.
  • Attributable — a named individual, not a shared account or a system user.
  • Time-stamped — including how long after the draft became available the signature was applied.
  • Exportable — a self-contained artifact a reviewer can verify without access to your vendor.

What to ask a vendor, and what the answers mean

Most AI scribe products are built to minimise the time between encounter and signed note. That is the feature they sell, and for many users it is the right trade. It does mean the review step is often the thinnest part of the product, and the questions below distinguish a tool that records review from one that merely allows it.

  • Does the system record who reviewed a note and when, as a stored record rather than a UI state?
  • Can it show the elapsed time between the draft being generated and the signature being applied?
  • If a note is edited after signing, is the change detectable from the record alone?
  • Can the review record be exported as a standalone document, without a login?
  • Does the record distinguish a note a person verified from one that went into the chart unreviewed?

Where this leaves AI documentation

The conclusion is not that AI-drafted notes are unsafe to bill. It is that the review step is the part carrying the regulatory weight, and that treating it as a formality converts a time saving into an evidentiary liability. A note drafted by software and genuinely reviewed by a clinician is a valid record, and the drafting tool is doing exactly what a tool should.

The failure mode to guard against is a review that exists in the workflow but not in the record. That is also the failure mode hardest to detect internally, because everything looks correct until something is challenged — and by then the interval between draft and signature is already in the log, and will be read by someone whose job is to find it.

One caution on scope. This describes documentation and billing requirements, which are federal and reasonably uniform. Malpractice standards are state law and vary, and nothing here is a substitute for advice from counsel or your compliance officer on your own arrangements.

Sources

  • 42 U.S.C. 1395l(e) — payment conditioned on the provider furnishing necessary information
  • 42 C.F.R. 482.24(c)(1) — medical record entries must be legible, complete, dated, timed, and authenticated
  • 42 C.F.R. 424.5(a)(6) — provider obligation to furnish information about services billed
  • CMS Medicare Program Integrity Manual, Pub. 100-08, Chapter 3 — verifying potential errors and taking corrective action
  • 45 C.F.R. 164.312(b) — audit controls for systems containing electronic protected health information

Verified 19 September 2026.

The regulatory information on this page is general background compiled from public primary sources, not legal or compliance advice. Requirements change and vary by jurisdiction and by court. Verify current rules with the relevant authority or your own counsel before relying on them.

Need transcription you can rely on?

Files are transcribed with AI, and you choose per upload whether a human reviewer checks the result. We confirm the price before work begins.