Skip to content
Back to CMS Watch
PolicyAugust 21, 2026·6 min read

Ambient AI scribes omit more than they invent, and an omitted chronic condition is a lost HCC

Two 2026 peer-reviewed evaluations point the same way: the dominant ambient AI scribe error is leaving content out, not making it up. One reported accidental omissions as the single most frequent error type; the other found the widest human-vs-AI quality gap was in how thorough the note was. For a risk-adjustment coder, a chronic condition dropped from the note is an HCC you never get to code.

AI documentationambient scriberisk adjustmentHCC recaptureV28
HCC Buddy

By the HCC Buddy Coding Team

Published August 21, 2026

A dictation microphone beside a chart folder with one file card pulled out, showing an ambient AI note that omits a condition.
The error a coder is positioned to catch in an AI note is the condition that quietly did not make it in.Image: HCC Buddy

Key Takeaways

  • A JMIR Medical Informatics pilot published April 17, 2026, in which 31 physicians drafted 7,545 notes with an ambient AI scribe, reported that among 356 reviewed notes accidental omissions were the most frequent error type (18%), ahead of hallucinations (11.5%) and accidental inclusions (9.3%), while 94.7% of notes were free of significant errors.
  • A vendor-neutral Annals of Internal Medicine evaluation (June 2026) scored 11 AI scribe tools against human notes on five standardized primary-care cases and found AI notes lower on all 10 modified PDQI-9 domains, with the largest deficits in being thorough, organized, and useful.
  • Read together, the two 2026 evaluations point the same way: the dominant residual ambient-AI note error is leaving content out, not fabricating it.
  • An omitted diagnosis is invisible to a downstream coder because there is no discrepancy in the record to catch, unlike a hallucinated diagnosis, which a coder can check for support and drop.
  • Under the CMS-HCC V28 model that pays in payment year 2026, a chronic condition dropped from a note goes uncoded, so its HCC and RAF weight are not recaptured for the year, a revenue gap rather than an audit finding.

Most ambient AI scribe notes are clean. The residual error that isn't clean tends to be the note leaving something out, not the note making something up, and that is the error a coder is positioned to catch. Two peer-reviewed evaluations published in 2026 landed on that same pattern from opposite directions.

The first is a pragmatic pilot in JMIR Medical Informatics (published April 17, 2026), in which 31 physicians used an ambient AI scribe to draft 7,545 clinic notes and reviewed a sample of them for four error types. The second is a vendor-neutral comparison in the Annals of Internal Medicine (June 2026), which scored AI-generated notes against human-written notes across five standardized primary-care cases. Neither study is about diagnosis codes. The risk-adjustment reading below is ours, not theirs.

What the two studies actually found

The JMIR pilot ran from July to August 2024. Physicians evaluated 356 (4.7%) of the AI-generated notes and rated each error's potential to harm a patient if left uncorrected. The headline is reassuring: the authors report that 94.7% of reviewed notes were free of significant errors.

The part that matters for a coder is the breakdown of the errors that did occur. Accidental omissions were the most frequent error type, in 18% of reviewed notes. Fabrication was rarer: hallucinations showed up in 11.5% of reviewed notes, and accidental inclusions, which carry real but misplaced content rather than invented content, in 9.3%; bias was rare at 1.1%. The authors' own conclusion is that "careful clinician review of notes remains imperative."

The Annals evaluation, funded by the Veterans Health Administration, took a different route to a similar place. Eleven AI scribe tools, 18 human note-takers, and 30 blinded raters worked from five audio-recorded standardized cases. Raters scored every note on the modified Physician Documentation Quality Instrument (PDQI-9), a 10-domain, 50-point scale. The authors report that human notes scored higher than AI notes across all five cases, with the largest single-case gap in acute low back pain (human 43.8 vs. AI 20.3). In the pooled domain analysis, AI scored lower on all 10 domains, and the three largest deficits were being thorough, organized, and useful.

Read together, one study says the most common ambient-AI error is omission, and the other says the biggest quality gap is thoroughness. Those are the same finding in two vocabularies.

Why omission is the coder's problem and not the physician's

A physician editing an AI draft is comparing it against the visit they just had. They know the patient's COPD came up, so if the draft skips it they can add it back. That is the correction loop both studies assume, and it mostly works.

A coder never had the visit. A coder has the note. If a condition was discussed and the note left it out, there is nothing in the record to react to, no flag, no discrepancy, no cue. A hallucinated diagnosis is at least visible: a coder checks it for support under the MEAT criteria, finds none, and drops it. An omitted diagnosis is invisible by definition. It is the one error class the downstream coder cannot catch by reading harder, because the thing to catch isn't on the page.

That inverts the usual worry. The compliance conversation about AI notes has centered on the tool adding an unsupported code, which is the failure mode behind the ambient-scribe risk-adjustment exposure and the capture-delta audit. Omission is the mirror image. For risk adjustment it costs money, and it does so silently.

For risk adjustment, a dropped chronic condition is a lost recapture

Under the CMS-HCC V28 model that pays in 2026, most chronic conditions have to be documented and coded every calendar year to keep contributing to the patient's risk score. A condition that was real, was discussed, and simply never made it into the note is not an audit problem. It is a condition that goes uncoded, so its HCC and its RAF weight fall off the risk score for the year. That is the recapture gap, and it is the opposite of the overcoding story: nobody gets a letter, the money just isn't there.

The conditions most exposed are exactly the ones a clinician mentions in passing on a stable patient: the long-standing diabetes that isn't today's chief complaint, the CKD noted from a lab trend, the COPD that is controlled and barely comes up. Those are the lines an ambient draft is likeliest to compress out, and each one carries an HCC.

The chronic conditions most likely to get dropped, and what they carry

Every code and HCC below was verified against the CMS-HCC V28 model for payment year 2026. Community weights are shown; the weight your plan realizes depends on the enrollee's segment.

Passing mention in the visitCodeV28 HCC (PY2026)Community RAF
Controlled COPD, not today's complaintJ44.9HCC 280, Chronic Obstructive Pulmonary Disease, Interstitial Lung Disorders, and Other Chronic Lung Disorders0.319
CKD stage 3b noted from a lab trendN18.32HCC 328, Chronic Kidney Disease, Moderate (Stage 3B)0.127
Long-standing type 2 diabetes, stableE11.9HCC 38, Diabetes with No, Glycemic, or Unspecified Complications0.166

None of these is an edge case. They are the everyday chronic load of a Medicare Advantage panel, and they are the conditions a recapture pass exists to catch. An ambient note that drops one is a recapture you have to find another way.

The completeness cross-check, since the studies stop at naming the problem

The QA habit for AI notes has mostly been "read the note against the visit." A coder can't do that. So build a check the coder can actually run, and run it against the two things a coder does have: the patient's own history and the note in front of them.

  • Reconcile against last year's captured HCCs. Pull the conditions that mapped to an HCC for this patient in the prior year. For each one, confirm it is either addressed in this year's note or has a documented reason it isn't (resolved, ruled out, no longer active). A chronic condition that was captured last year and is simply absent this year, with no explanation, is the omission to chase down before you move on.
  • Reconcile against the problem list. Where the active problem list carries a chronic condition and the AI-drafted note is silent on it, that is a query, not a code. Do not code from the problem list alone, and do not treat the note's silence as a resolution.
  • Send it back to the provider. The fix for a suspected omission is a provider query, because the coder can flag that a captured chronic condition went unaddressed but cannot document that it was addressed. Reconciling is the coder's job; re-documenting is the provider's.

This is the same discipline as the QA pass every AI-drafted note still needs, narrowed to the one failure mode these two studies say is the most common: the condition that isn't there.

What the studies did not test

Keep the limits straight, because neither study measured coding at all. The JMIR work is a single-health-system pilot in which physicians self-evaluated a 4.7% sample of their own tool's notes, so the omission rate is an estimate from one setting, not a benchmark. The Annals evaluation used simulated standardized cases, and the authors note the human notes were not produced under real-world time pressure, which flatters the human comparison. Neither study reported an ICD-10-CM code, an HCC, or a RAF figure. The connection from "AI notes omit content" to "an HCC goes uncaptured" is a coding inference from the guidelines and the V28 model, not a measured result, and it is worth stating plainly to anyone who asks.

Where a human coder is still required

An omission has no signature. There is no confidence score on a note that flags the sentence that should have been there, and no fluency metric that penalizes the tool for being too brief about a stable chronic condition. The only thing that surfaces a dropped diagnosis is a person reconciling the note against what the patient is already known to carry.

That reconciliation is also the work that decides whether the risk score is complete and whether the documentation behind each condition will hold when someone asks why the code is, or isn't, on the claim. A tool that quietly leaves out the COPD hasn't made an error the tool can see. It has made one only a coder can see, and only if the coder goes looking.

What coders should do now

  1. 1Reconcile the AI-drafted note against the patient's prior-year captured HCCs. For each chronic condition captured last year, confirm this year's note either addresses it or documents why it no longer applies; an unexplained absence is the omission to chase.
  2. 2Cross-check the note against the active problem list. Where the problem list carries a chronic condition the note is silent on, treat it as a provider query, not a code, and never code from the problem list alone.
  3. 3Prioritize the stable chronic conditions that get mentioned in passing, such as controlled COPD (J44.9), CKD noted from a lab trend (N18.32), or long-standing type 2 diabetes (E11.9), since those are the lines an ambient draft is likeliest to compress out and each carries a V28 HCC.
  4. 4Send a suspected omission back to the provider with a [provider query](/blog/provider-query-templates); the coder can flag that a captured condition went unaddressed but cannot document that it was addressed.
  5. 5Run the recovered condition through your encoder to confirm it still maps to an HCC under the current V28 model before it goes on the claim.

Frequently Asked Questions

What is the most common error in ambient AI scribe notes?

A JMIR Medical Informatics pilot published April 17, 2026 reported that accidental omissions were the most frequent error type among reviewed notes (18%), ahead of hallucinations (11.5%) and accidental inclusions (9.3%). The study also reported that 94.7% of reviewed notes were free of significant errors, so most notes were clean and omission was the leading error among those that were not. It was a single-health-system pilot in which physicians self-evaluated a 4.7% sample.

Why is an omission in an AI note worse for a coder than a hallucination?

A hallucinated diagnosis is visible: a coder checks it for supporting documentation and, finding none, drops it. An omitted diagnosis leaves no trace in the record, so there is no discrepancy for the coder to catch by reading more carefully. The physician who had the visit can add a missing condition back; the coder, who only has the note, cannot see that anything is missing.

How does a dropped chronic condition affect a Medicare Advantage risk score?

Under the CMS-HCC V28 model used for payment year 2026, most chronic conditions must be documented and coded each calendar year to keep contributing to the patient's risk score. A condition that was discussed but left out of the note goes uncoded, so its HCC and RAF weight are not recaptured for that year. That is a revenue gap rather than an audit finding, because nothing unsupported was submitted, the supported code simply never was.

Did these studies measure diagnosis coding or HCC capture?

No. The JMIR pilot classified note errors (omissions, hallucinations, inclusions, bias) and the Annals evaluation scored documentation quality on the modified PDQI-9. Neither reported an ICD-10-CM code, an HCC, or a RAF figure. The risk-adjustment consequence is a coding inference drawn from the guidelines and the V28 model, not a measured result of either study.

How can a coder catch a condition an AI note left out?

Reconcile the note against the two things a coder actually has: the patient's prior-year captured HCCs and the active problem list. A chronic condition that mapped to an HCC last year, or that sits on the problem list, but is absent from this year's AI-drafted note with no explanation, is the omission to query. The coder flags the gap; the provider documents whether the condition was addressed.

Related topics:AI documentationambient scriberisk adjustmentHCC recaptureV28
HCC Buddy

HCC Buddy Coding Team

Editorial

Every HCC Buddy news article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.

Get CMS Updates in Your Inbox

RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.