Skip to content
Back to CMS Watch
PolicySeptember 8, 2026·5 min read

AI scribe drafts can read more certain than the visit was, a 2026 study finds

Clinicians revising ambient-AI note drafts more often added hedging than removed it, a 2026 preprint reports, so AI text can overstate diagnostic certainty. In outpatient risk adjustment a probable or suspected diagnosis cannot be coded, so certainty in an AI-drafted note is not the same as a confirmed one.

AI documentationambient scribeuncertain diagnosisrisk adjustmentRADV
HCC Buddy

By the HCC Buddy Coding Team

Published September 8, 2026

Ambient voice-capture microphone by a laptop with a blurred clinical note, illustrating how an AI scribe draft can overstate certainty.
An ambient AI scribe drafts the note. The coder still confirms whether a diagnosis is as certain as the draft makes it look.Illustration: HCC Buddy

Key Takeaways

  • In a 2026 arXiv preprint analyzing 62,811 paired note sections from outpatient ambient-AI documentation platforms, University of California, Irvine researchers reported that clinicians introduced hedging language into previously unhedged text more often than they removed it.
  • Where clinicians rewrote hedging language, the edits shifted toward greater uncertainty rather than greater certainty, which the authors read as clinicians recalibrating AI drafts they judged too definitive; the pattern held across both vendors studied.
  • Under the FY2026 ICD-10-CM Official Guidelines Section IV.H, outpatient and physician-office coders may not code a diagnosis documented as probable, suspected, questionable, rule out, or working diagnosis, and must code the sign, symptom, or reason for the visit instead.
  • The inpatient rule (Section II.H) is the opposite, coding an uncertain diagnosis as if established, so the setting determines whether hedged language can be coded at all.
  • The study measured clinician editing behavior, not coding outcomes; the risk that an uncorrected over-certain draft reaches a claim is an inference from the finding, not a rate the study reports.

A 2026 preprint posted to arXiv put a number on something coders have half-suspected about AI-drafted notes. When clinicians rewrote the drafts an ambient-AI scribe produced, they more often added words that soften a statement than removed them. Researchers at the University of California, Irvine compared each AI draft against the clinician's signed note across 62,811 paired note sections, drawn from two commercial ambient-AI platforms piloted in outpatient clinics. On balance the edits leaned toward more uncertainty, which the authors read as clinicians pulling back drafts they judged too definitive. The plain reading for a coder: the draft tends to sound surer than the clinician was.

What the study measured

The team compared the AI draft and the signed note section by section across up to four parts of a note (history, exam, assessment and plan, and results), then counted how a curated list of 536 hedging terms moved between draft and final text. The analysis covered two vendors, anonymized as Vendor A and Vendor B, and both primary care and specialty encounters. This is a preprint, so it has not yet completed peer review, and every figure below is the authors' reported result.

Which way the edits went

The study's net measure of edit direction points toward uncertainty. Among the sections where a clinician made a hedging-related directional edit, more shifted the wording toward less certainty than toward more, and the authors report that net tendency as statistically significant. It applied to a minority of edits, though. Most note sections carried no hedging-related change at all, so this is a signal in the margins, not a rewrite of every note.

The single most common word substitution the study logged was "likely" softened to "possibly." The paper catalogs more than a thousand such substitutions, in both directions and sorted by frequency, so read that list as the texture of the editing rather than proof that every change ran one way. The direction comes from the net-score analysis, not the word list.

The authors' own framing is the useful part for a coder. They write that tentative phrasing should reflect whether a diagnosis is confirmed or still under consideration, and whether supporting evidence from labs, imaging, or pathology is present. That's a coding standard stated in clinical language. They describe the pattern as likely a cross-vendor issue rather than one platform's quirk, since both platforms they studied trended the same way.

Why an over-certain draft is a coding problem

Read the finding against the FY2026 ICD-10-CM Official Guidelines. The rule for the professional and physician-office record is explicit: you do not code a diagnosis documented as probable, suspected, questionable, rule out, compatible with, consistent with, or working diagnosis. You code the sign, the symptom, or the reason for the visit instead. The inpatient rule is the opposite, and that split is what makes an over-certain draft a live risk in the outpatient and physician-office setting where much risk-adjustment coding happens.

SettingAn uncertain diagnosis ("probable," "suspected," "rule out")Guideline
Inpatient discharge (short-term acute, long-term care, psychiatric)Code it as if establishedSection II.H
Outpatient and physician officeDo not code it; code the sign, symptom, or reason for the visitSection IV.H

So the certainty word in the note is load-bearing. "Possible sepsis" in a clinic note does not support a sepsis code; you fall back to the documented signs. If an AI draft renders that same clinical picture as "sepsis," flat, and it is signed without the qualifier the clinician would have added, the note now reads like a confirmed diagnosis the encounter never established.

Where risk adjustment raises the stakes

Much of risk-adjustment coding happens in the outpatient and physician-office record, which follows the outpatient rule. A condition that maps to an HCC has to be supported by the record, and RADV review exists to confirm exactly that. A diagnosis that reads as confirmed but was never clinically established is not supported documentation, and unsupported HCCs are what a RADV audit recovers.

It's the mirror image of the scribe risk coders already watch, where the draft drops a chronic diagnosis and costs a recapture. One failure under-documents the chart. This one over-states it, and both land on the coder to catch.

That's the gap between how the study frames its finding (a clinician editing habit) and why it lands on a coder's desk. The clinician who catches the over-certain draft and softens it has closed the gap. The coder is the backstop for the note where that edit didn't happen.

What the study did not show

Be precise about what this is and isn't. The study measured notes clinicians did revise, and the shift toward uncertainty is evidence that clinicians treated the drafts as too definitive. It didn't measure coding outcomes, it didn't count how many over-certain drafts were signed without correction, and it didn't break the finding out by the assessment section specifically, where a diagnosis actually lives. The coder-facing risk (an uncorrected over-certain note reaching the claim) is a reasonable inference from the pattern, not a number the study reports. Treat it as a reason to check, not a measured error rate. The work also carries the usual preprint caveats: a single health system, no patient or clinician characteristics, and no adjustment for which vendor was used in which specialty.

Where the coder still owns the code

The study's own recommendation (make tentative phrasing match whether the diagnosis is actually established) is a coder's reflex, not a software feature. When you code from an ambient-AI or scribe-generated outpatient note, the certainty of a diagnosis is a claim to verify against the rest of the record, not a fact to inherit. Does the assessment's confidence match the labs, the imaging, the exam? Is there MEAT behind the condition, or just a confident sentence? When the certainty outruns the evidence, that is a provider query, the same call a coder has always made, now aimed at a draft a machine wrote first. That judgment is the part no ambient scribe in this study could do on its own.

What coders should do now

  1. 1When you code a diagnosis from an ambient-AI or scribe-generated outpatient note, treat its stated certainty as a claim you confirm against the record before you rely on it. Check that the assessment's confidence matches the labs, imaging, and exam findings in the same note before you assign a risk-adjusted code.
  2. 2Apply the outpatient rule without exception: a diagnosis documented as probable, suspected, rule out, or working diagnosis is not codeable in the physician-office record, per FY2026 Guideline IV.H. Code the sign or symptom instead, no matter how confident the AI draft reads.
  3. 3When the note's certainty outruns the evidence, query the provider to confirm or soften the diagnosis before it goes on the claim, rather than coding the confident sentence as written.
  4. 4Line up the [MEAT](/meat-criteria) for any risk-adjusted diagnosis you pull from an AI draft. A condition asserted with no monitoring, evaluation, assessment, or treatment behind it has cosmetic certainty, not support.
  5. 5For your highest-RAF conditions, spot-check a sample of AI-drafted charts for certainty that was never clinically established, and line up the [evidence](/evidence) the way a RADV reviewer would before it is requested.

Frequently Asked Questions

Do AI scribe notes make diagnoses sound more certain than they are?

A 2026 arXiv preprint found that when clinicians edited ambient-AI note drafts, they added hedging language more often than they removed it and shifted wording toward greater uncertainty, which the authors read as clinicians correcting drafts they judged too definitive. It is a study of editing behavior, so treat it as a signal that AI drafts can overstate certainty, not a measured error rate.

Can you code a probable or suspected diagnosis in outpatient risk adjustment?

No. Under the FY2026 ICD-10-CM Official Guidelines Section IV.H, a diagnosis documented as probable, suspected, questionable, rule out, compatible with, consistent with, or working diagnosis is not coded in the outpatient or physician-office record. You code the sign, symptom, or reason for the visit instead. Because a large share of risk-adjustment coding happens in that professional record, hedged language there usually means no HCC.

Why does an over-certain AI note matter for RADV?

A diagnosis that maps to an HCC has to be supported by the medical record, and RADV review exists to confirm it. If an AI draft renders a still-uncertain clinical picture as a flat, confirmed diagnosis and it's signed without correction, the note can read as established when the encounter didn't establish it, and an unsupported HCC is what a RADV audit recovers.

Does this study prove AI notes cause miscoding?

No. The study measured how clinicians edit AI drafts, not coding accuracy or claims. It didn't count how many over-certain drafts were signed uncorrected, and it didn't isolate the assessment section. The coding risk is a reasonable inference from the finding that drafts skew definitive. It's a reason for coders to verify certainty, not a demonstrated coding outcome.

What should a coder check on an AI-drafted note before coding a diagnosis?

Confirm the diagnosis's stated certainty matches the supporting evidence in the note (labs, imaging, exam), apply the outpatient uncertain-diagnosis rule so hedged conditions fall back to signs and symptoms, verify MEAT for any risk-adjusted condition, and query the provider when the note reads more certain than the workup supports.

Related topics:AI documentationambient scribeuncertain diagnosisrisk adjustmentRADV
HCC Buddy

HCC Buddy Coding Team

Editorial

Every HCC Buddy news article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.

Get CMS Updates in Your Inbox

RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.