Skip to content
Back to CMS Watch
PolicyJuly 20, 2026·5 min read

GAO: AI coding tools may capture more diagnoses, and few studies have checked their accuracy

GAO says AI tools that suggest codes for coder review are already in widespread use, that newer tools can assign codes on their own, and that they could capture more diagnoses and services than a non-AI workflow. Here's how to measure that over-capture on your own charts.

AI documentationmedical codingGAOdocumentation integrityaudit risk
Jess P., CPC

Reviewed by Jess P., CPC

Published July 20, 2026

A magnifier resting on a printed report beside stacked chart folders, illustrating scrutiny of AI medical coding accuracy
GAO's own conclusion: the accuracy of AI note and coding tools may be difficult to verify.Image: HCC Buddy

Key Takeaways

  • GAO published Science & Tech Spotlight GAO-26-109116, "AI for Medical Notes and Coding," on July 16, 2026, concluding that the accuracy of AI note and coding tools may be difficult to verify and that effects on health care spending are uncertain.
  • GAO states that AI tools suggesting codes for coder review are already in widespread use, and that newer generative and agentic tools able to assign codes autonomously "could make human coders faster or replace them entirely."
  • The report's "more than 95 percent accuracy" figure is attributed by GAO to a single unnamed AI coding software developer describing its own deployment at one five-hospital health system, not to an independent evaluation.
  • GAO warns AI scribing and medical coding tools could increase health care costs if they capture more diagnoses and services during visits than non-AI approaches, and that inaccuracies may cause over- or under-reimbursement.
  • GAO cites an American Medical Association finding that the share of surveyed clinicians using AI for clinical documentation or medical coding rose from 21 percent to 28 percent between 2024 and 2026.

A federal watchdog now says in writing that AI coding tools may pull more diagnoses out of an encounter than a coder would, and that hardly anyone outside the vendors has measured how accurate they are. Every extra code still needs documentation behind it, and the ones that don't are the population an auditor samples.

That's the Government Accountability Office's Science & Tech Spotlight "AI for Medical Notes and Coding" (GAO-26-109116), published July 16, 2026. Two pages, no GAO-collected sample, and GAO's own footer says it "is not an audit product." It matters anyway, because there are few independent studies evaluating whether these tools are accurate, and the one accuracy figure GAO relays came from a vendor describing its own product.

GAO's summary line is the one to keep: "The accuracy of these tools may be difficult to verify, and the overall effects on health care spending are uncertain."

What GAO said about human coders

GAO doesn't hedge on where this is going. In its own voice:

"AI tools that analyze patient records and suggest codes for review by medical coders are already in widespread use. New AI tools may use generative and agentic AI technologies to review patient records and assign codes autonomously. This capability could make human coders faster or replace them entirely."

That's a federal agency putting the displacement question in a public document. Note the split GAO draws and most coverage won't: suggest-for-review tools are the ones already everywhere. Autonomous assignment is the new thing.

The rest of the Spotlight is why it isn't settled. GAO says there are "few independent studies evaluating the accuracy of these tools," and that some tools may not store recordings and transcripts, "which may limit the extent to which health care providers can conduct independent assessments."

The 95% number is a vendor's own, and GAO says so

The figure that'll get pulled out of this report and pasted into slide decks reads, in GAO's text:

"According to one AI medical coding software developer, when its software was deployed at a health system with five hospitals, it generated medical codes with more than 95 percent accuracy, and the system's emergency departments reduced annual coding costs by more than $1 million."

One unnamed developer. One health system. Reported by the developer. GAO relays it and, in the same document, concludes accuracy is difficult to verify because independent studies are few. Anyone quoting the 95% without the sentence that follows it is quoting a marketing claim with a government seal borrowed onto it.

Who actually said what

ClaimWhose claim it isWhat it rests on
AI tools that suggest codes for coder review are already in widespread useGAOGAO's own statement
Newer generative and agentic tools may assign codes autonomously, which "could make human coders faster or replace them entirely"GAOGAO's own statement
Accuracy may be difficult to verify; few independent studies existGAOGAO's own statement
More than 95% coding accuracy; over $1M in annual ED coding cost reductionOne unnamed AI coding software developerVendor self-report, one five-hospital deployment, relayed by GAO
Clinician use of AI for documentation or coding rose from 21% to 28% between 2024 and 2026American Medical AssociationAMA survey of responding clinicians, cited by GAO
Documentation time reduced 20 percent, or two minutes per appointment, using AI scribes"One study," cited by GAOSingle study, not named in the Spotlight; a scribe result, not a coding-tool result
U.S. clinicians average a 57-hour workweek including 7 hours of administrative workGAOGAO's own statement; the Spotlight names no source for it

The challenge that actually lands on your desk

Of GAO's four challenges, one is a coding problem rather than a policy problem. GAO writes that these tools "could increase health care costs if they capture more diagnoses and services rendered during visits than non-AI approaches," and that inaccuracies "may result in patient harm or over- or under-reimbursement from insurers to providers."

Read that from a coder's chair and it stops being an abstraction. It is the same problem as an AI-drafted note that still needs coder QA, one step further down the pipe: now the tool isn't just writing the narrative, it's proposing the codes. If a tool systematically pulls more diagnoses out of the same encounter than a coder would, the extra codes aren't free. Each one either has documentation behind it or it doesn't, and the ones that don't are exactly what an auditor samples. "The tool suggested it" has never been a defense.

GAO's four challenges, and what you do about each

GAO challengeWhat it means at the deskWhat you do about it
Difficulties verifying accuracyVendor accuracy figures are self-reported and not independently reproducedMeasure it on your own charts instead of accepting the deck. See the capture-delta audit below
Limited access due to costsSmall practices get the tool later, or a cheaper one with weaker review featuresConfirm the tool shows you the source note text behind every suggested code before you buy
Reimbursement and cost implicationsTools may capture more diagnoses and services than a non-AI workflowAudit the delta, not the total. The shared codes were already yours; the added ones are the new exposure
Data privacy and patient consentData retention practices vary across AI scribe vendors, and patients may not always be informed that recordings are occurringAsk what's retained and for how long, because it decides whether you can ever re-audit an encounter

The capture-delta audit

GAO names the over-capture risk and stops. It's a watchdog Spotlight, not a procedure. Here's the procedure.

Pull a sample of encounters the tool has already coded. Have a credentialed coder assign codes from the same record, blind to the tool's output. Then diff the two sets in both directions.

Every code the tool assigned and the coder didn't goes into one of three piles:

  • Supported: the documentation's there, the coder just missed it. That's the tool earning its keep.
  • Query pile: the dx is in the note, but there's no monitoring, evaluation, assessment or treatment behind it at that DOS. Not a code yet. That's a query.
  • Not supported in the documentation: nothing in the record backs it. That's the pile that turns into a repayment.

Then run the reverse column: every code the coder assigned and the tool didn't. GAO names under-reimbursement in the same breath as over-, and a tool that quietly drops a supported chronic condition costs you a recapture nobody flagged.

How those piles split is the only accuracy number for your shop that nobody has to take a vendor's word for. Run it again after any model or version change, because the answer isn't stable across updates.

This is the chart-level version of the claims-level check in coding intensity is climbing where AI documentation landed. That one compares documented severity against documented treatment across a population. This one compares two coders on one chart.

Where a human coder is still required

All four of GAO's challenges come back to a question a tool can't answer for you: is this diagnosis actually addressed at this encounter, or is it just in the note? That's a read of what the provider documented against what the code set demands, and if the answer is no, there's a repayment on the other side of it.

GAO says policymakers need more information about the performance of these tools to determine the appropriate level of oversight. That information doesn't exist yet. Until it does, the only accuracy number your shop has is the one you measured yourself.

What coders should do now

  1. 1Run the capture-delta audit on a live sample and log the split. The added codes that need a query first, plus the ones with nothing behind them, are your exposure. That number, not the vendor's percentage, is the one to report upward.
  2. 2Diff in both directions. Codes the tool dropped that the coder supported are a recapture problem, and GAO names under-reimbursement alongside over-reimbursement.
  3. 3Stop repeating vendor accuracy percentages in internal decks unless the sentence names who measured it. GAO's own report shows how far a self-reported number travels once it's quoted without its attribution.
  4. 4Ask your documentation vendor two questions in writing: what source text sits behind each suggested code, and how long recordings and transcripts are retained. GAO flags non-retention as a barrier to independent assessment, and it decides whether you can ever re-audit the encounter.
  5. 5Treat every AI-added chronic diagnosis as a query question before it becomes a code. Check the encounter against the [MEAT criteria](/meat-criteria) rather than the tool's confidence score.

Frequently Asked Questions

Did GAO say AI will replace medical coders?

GAO stated that new generative and agentic AI tools able to review records and assign codes autonomously "could make human coders faster or replace them entirely." That is framed as a capability and a possibility, not a prediction or a finding. In the same report GAO concludes there are few independent studies evaluating whether these tools are accurate enough to rely on.

Is the 95% AI coding accuracy figure from GAO?

No. GAO attributes it to one AI medical coding software developer describing its own software deployed at a health system with five hospitals. It is a vendor self-report relayed by GAO, not a GAO measurement or an independent evaluation, and GAO separately concludes that accuracy may be difficult to verify because few independent studies exist.

What did GAO identify as the risks of AI medical coding tools?

GAO identified four challenges: difficulties verifying accuracy, limited access due to costs, reimbursement and cost implications, and data privacy and patient consent. On cost, GAO warns the tools could increase health care spending if they capture more diagnoses and services during visits than non-AI approaches.

How do I audit an AI coding tool's accuracy?

Recode a sample of tool-coded encounters blind, then look only at the codes the tool added and the coder did not. The ones with no documentation behind them, plus the ones that would need a provider query first, are your exposure. Check the reverse direction too, since a tool that drops a supported condition costs a recapture. Unlike a vendor accuracy figure, that number came off your own charts.

How many clinicians are using AI for documentation or coding?

GAO cites an American Medical Association finding that between 2024 and 2026 the share of surveyed clinicians using AI tools to assist with clinical documentation or medical coding increased from 21 percent to 28 percent. That is a survey estimate of clinicians who responded, not a measured national adoption rate.

Related topics:AI documentationmedical codingGAOdocumentation integrityaudit risk
Jess P., CPC

Jess P., CPC

Certified Professional Coder

Jess reviews HCC Buddy editorial content for accuracy against the current CMS-HCC model and the active FY ICD-10-CM tabular release.

Get CMS Updates in Your Inbox

RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.