Ambient AI scribes omit more than they invent, and an omitted chronic condition is a lost HCC
Two 2026 peer-reviewed evaluations point the same way: the dominant ambient AI scribe error is leaving content out, not making it up. One reported accidental omissions as the single most frequent error type; the other found the widest human-vs-AI quality gap was in how thorough the note was. For a risk-adjustment coder, a chronic condition dropped from the note is an HCC you never get to code.
By the HCC Buddy Coding Team
Published August 21, 2026

Key Takeaways
- →A JMIR Medical Informatics pilot published April 17, 2026, in which 31 physicians drafted 7,545 notes with an ambient AI scribe, reported that among 356 reviewed notes accidental omissions were the most frequent error type (18%), ahead of hallucinations (11.5%) and accidental inclusions (9.3%), while 94.7% of notes were free of significant errors.
- →A vendor-neutral Annals of Internal Medicine evaluation (June 2026) scored 11 AI scribe tools against human notes on five standardized primary-care cases and found AI notes lower on all 10 modified PDQI-9 domains, with the largest deficits in being thorough, organized, and useful.
- →Read together, the two 2026 evaluations point the same way: the dominant residual ambient-AI note error is leaving content out, not fabricating it.
- →An omitted diagnosis is invisible to a downstream coder because there is no discrepancy in the record to catch, unlike a hallucinated diagnosis, which a coder can check for support and drop.
- →Under the CMS-HCC V28 model that pays in payment year 2026, a chronic condition dropped from a note goes uncoded, so its HCC and RAF weight are not recaptured for the year, a revenue gap rather than an audit finding.
Most ambient AI scribe notes are clean. The residual error that isn't clean tends to be the note leaving something out, not the note making something up, and that is the error a coder is positioned to catch. Two peer-reviewed evaluations published in 2026 landed on that same pattern from opposite directions.
The first is a pragmatic pilot in JMIR Medical Informatics (published April 17, 2026), in which 31 physicians used an ambient AI scribe to draft 7,545 clinic notes and reviewed a sample of them for four error types. The second is a vendor-neutral comparison in the Annals of Internal Medicine (June 2026), which scored AI-generated notes against human-written notes across five standardized primary-care cases. Neither study is about diagnosis codes. The risk-adjustment reading below is ours, not theirs.
What the two studies actually found
The JMIR pilot ran from July to August 2024. Physicians evaluated 356 (4.7%) of the AI-generated notes and rated each error's potential to harm a patient if left uncorrected. The headline is reassuring: the authors report that 94.7% of reviewed notes were free of significant errors.
The part that matters for a coder is the breakdown of the errors that did occur. Accidental omissions were the most frequent error type, in 18% of reviewed notes. Fabrication was rarer: hallucinations showed up in 11.5% of reviewed notes, and accidental inclusions, which carry real but misplaced content rather than invented content, in 9.3%; bias was rare at 1.1%. The authors' own conclusion is that "careful clinician review of notes remains imperative."
The Annals evaluation, funded by the Veterans Health Administration, took a different route to a similar place. Eleven AI scribe tools, 18 human note-takers, and 30 blinded raters worked from five audio-recorded standardized cases. Raters scored every note on the modified Physician Documentation Quality Instrument (PDQI-9), a 10-domain, 50-point scale. The authors report that human notes scored higher than AI notes across all five cases, with the largest single-case gap in acute low back pain (human 43.8 vs. AI 20.3). In the pooled domain analysis, AI scored lower on all 10 domains, and the three largest deficits were being thorough, organized, and useful.
Read together, one study says the most common ambient-AI error is omission, and the other says the biggest quality gap is thoroughness. Those are the same finding in two vocabularies.
Why omission is the coder's problem and not the physician's
A physician editing an AI draft is comparing it against the visit they just had. They know the patient's COPD came up, so if the draft skips it they can add it back. That is the correction loop both studies assume, and it mostly works.
A coder never had the visit. A coder has the note. If a condition was discussed and the note left it out, there is nothing in the record to react to, no flag, no discrepancy, no cue. A hallucinated diagnosis is at least visible: a coder checks it for support under the MEAT criteria, finds none, and drops it. An omitted diagnosis is invisible by definition. It is the one error class the downstream coder cannot catch by reading harder, because the thing to catch isn't on the page.
That inverts the usual worry. The compliance conversation about AI notes has centered on the tool adding an unsupported code, which is the failure mode behind the ambient-scribe risk-adjustment exposure and the capture-delta audit. Omission is the mirror image. For risk adjustment it costs money, and it does so silently.
For risk adjustment, a dropped chronic condition is a lost recapture
Under the CMS-HCC V28 model that pays in 2026, most chronic conditions have to be documented and coded every calendar year to keep contributing to the patient's risk score. A condition that was real, was discussed, and simply never made it into the note is not an audit problem. It is a condition that goes uncoded, so its HCC and its RAF weight fall off the risk score for the year. That is the recapture gap, and it is the opposite of the overcoding story: nobody gets a letter, the money just isn't there.
The conditions most exposed are exactly the ones a clinician mentions in passing on a stable patient: the long-standing diabetes that isn't today's chief complaint, the CKD noted from a lab trend, the COPD that is controlled and barely comes up. Those are the lines an ambient draft is likeliest to compress out, and each one carries an HCC.
The chronic conditions most likely to get dropped, and what they carry
Every code and HCC below was verified against the CMS-HCC V28 model for payment year 2026. Community weights are shown; the weight your plan realizes depends on the enrollee's segment.
| Passing mention in the visit | Code | V28 HCC (PY2026) | Community RAF |
|---|---|---|---|
| Controlled COPD, not today's complaint | J44.9 | HCC 280, Chronic Obstructive Pulmonary Disease, Interstitial Lung Disorders, and Other Chronic Lung Disorders | 0.319 |
| CKD stage 3b noted from a lab trend | N18.32 | HCC 328, Chronic Kidney Disease, Moderate (Stage 3B) | 0.127 |
| Long-standing type 2 diabetes, stable | E11.9 | HCC 38, Diabetes with No, Glycemic, or Unspecified Complications | 0.166 |
None of these is an edge case. They are the everyday chronic load of a Medicare Advantage panel, and they are the conditions a recapture pass exists to catch. An ambient note that drops one is a recapture you have to find another way.
The completeness cross-check, since the studies stop at naming the problem
The QA habit for AI notes has mostly been "read the note against the visit." A coder can't do that. So build a check the coder can actually run, and run it against the two things a coder does have: the patient's own history and the note in front of them.
- Reconcile against last year's captured HCCs. Pull the conditions that mapped to an HCC for this patient in the prior year. For each one, confirm it is either addressed in this year's note or has a documented reason it isn't (resolved, ruled out, no longer active). A chronic condition that was captured last year and is simply absent this year, with no explanation, is the omission to chase down before you move on.
- Reconcile against the problem list. Where the active problem list carries a chronic condition and the AI-drafted note is silent on it, that is a query, not a code. Do not code from the problem list alone, and do not treat the note's silence as a resolution.
- Send it back to the provider. The fix for a suspected omission is a provider query, because the coder can flag that a captured chronic condition went unaddressed but cannot document that it was addressed. Reconciling is the coder's job; re-documenting is the provider's.
This is the same discipline as the QA pass every AI-drafted note still needs, narrowed to the one failure mode these two studies say is the most common: the condition that isn't there.
What the studies did not test
Keep the limits straight, because neither study measured coding at all. The JMIR work is a single-health-system pilot in which physicians self-evaluated a 4.7% sample of their own tool's notes, so the omission rate is an estimate from one setting, not a benchmark. The Annals evaluation used simulated standardized cases, and the authors note the human notes were not produced under real-world time pressure, which flatters the human comparison. Neither study reported an ICD-10-CM code, an HCC, or a RAF figure. The connection from "AI notes omit content" to "an HCC goes uncaptured" is a coding inference from the guidelines and the V28 model, not a measured result, and it is worth stating plainly to anyone who asks.
Where a human coder is still required
An omission has no signature. There is no confidence score on a note that flags the sentence that should have been there, and no fluency metric that penalizes the tool for being too brief about a stable chronic condition. The only thing that surfaces a dropped diagnosis is a person reconciling the note against what the patient is already known to carry.
That reconciliation is also the work that decides whether the risk score is complete and whether the documentation behind each condition will hold when someone asks why the code is, or isn't, on the claim. A tool that quietly leaves out the COPD hasn't made an error the tool can see. It has made one only a coder can see, and only if the coder goes looking.
What coders should do now
- 1Reconcile the AI-drafted note against the patient's prior-year captured HCCs. For each chronic condition captured last year, confirm this year's note either addresses it or documents why it no longer applies; an unexplained absence is the omission to chase.
- 2Cross-check the note against the active problem list. Where the problem list carries a chronic condition the note is silent on, treat it as a provider query, not a code, and never code from the problem list alone.
- 3Prioritize the stable chronic conditions that get mentioned in passing, such as controlled COPD (J44.9), CKD noted from a lab trend (N18.32), or long-standing type 2 diabetes (E11.9), since those are the lines an ambient draft is likeliest to compress out and each carries a V28 HCC.
- 4Send a suspected omission back to the provider with a [provider query](/blog/provider-query-templates); the coder can flag that a captured condition went unaddressed but cannot document that it was addressed.
- 5Run the recovered condition through your encoder to confirm it still maps to an HCC under the current V28 model before it goes on the claim.
Frequently Asked Questions
What is the most common error in ambient AI scribe notes?
A JMIR Medical Informatics pilot published April 17, 2026 reported that accidental omissions were the most frequent error type among reviewed notes (18%), ahead of hallucinations (11.5%) and accidental inclusions (9.3%). The study also reported that 94.7% of reviewed notes were free of significant errors, so most notes were clean and omission was the leading error among those that were not. It was a single-health-system pilot in which physicians self-evaluated a 4.7% sample.
Why is an omission in an AI note worse for a coder than a hallucination?
A hallucinated diagnosis is visible: a coder checks it for supporting documentation and, finding none, drops it. An omitted diagnosis leaves no trace in the record, so there is no discrepancy for the coder to catch by reading more carefully. The physician who had the visit can add a missing condition back; the coder, who only has the note, cannot see that anything is missing.
How does a dropped chronic condition affect a Medicare Advantage risk score?
Under the CMS-HCC V28 model used for payment year 2026, most chronic conditions must be documented and coded each calendar year to keep contributing to the patient's risk score. A condition that was discussed but left out of the note goes uncoded, so its HCC and RAF weight are not recaptured for that year. That is a revenue gap rather than an audit finding, because nothing unsupported was submitted, the supported code simply never was.
Did these studies measure diagnosis coding or HCC capture?
No. The JMIR pilot classified note errors (omissions, hallucinations, inclusions, bias) and the Annals evaluation scored documentation quality on the modified PDQI-9. Neither reported an ICD-10-CM code, an HCC, or a RAF figure. The risk-adjustment consequence is a coding inference drawn from the guidelines and the V28 model, not a measured result of either study.
How can a coder catch a condition an AI note left out?
Reconcile the note against the two things a coder actually has: the patient's prior-year captured HCCs and the active problem list. A chronic condition that mapped to an HCC last year, or that sits on the problem list, but is absent from this year's AI-drafted note with no explanation, is the omission to query. The coder flags the gap; the provider documents whether the condition was addressed.
Sources
- Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study (JMIR Medical Informatics, DOI 10.2196/86474) — JMIR Medical Informatics, Apr 17, 2026
- Rapid Evaluation of Artificial Intelligence Technology Used for Ambient Dictation in Primary Care: Comparing the Quality of Documentation of AI-Generated and Human-Produced Clinical Notes (Annals of Internal Medicine, DOI 10.7326/ANNALS-25-02772) — Annals of Internal Medicine, Jun 1, 2026
- CMS-HCC Risk Adjustment Model, Version 28 (2024 Model Software / PY2026 factors) — CMS, Jan 1, 2026
Related Tools
RAF Calculator
See what each recaptured chronic condition is worth to the risk score, so you know which omissions cost the most to miss.
MEAT Criteria Guide
Check what a recovered diagnosis needs to be Monitored, Evaluated, Assessed, or Treated before you code it back.
Evidence Builder
Line up the documentation behind a condition so it holds when someone asks why the code is, or isn't, on the claim.
HCC Buddy Coding Team
Editorial
Every HCC Buddy news article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.
Get CMS Updates in Your Inbox
RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.
More from CMS Watch
CPT 2027 sharpens the AI taxonomy. How a tool's role is classified decides how it's coded.
August 14, 2026
PolicySolventum plans to separate its coding and CDI software unit, home to 360 Encompass
August 10, 2026
PolicyNew state laws bar AI-only claim downcoding: Indiana live July 1, Illinois signed July 10
August 5, 2026

