OIG's first MA compliance guidance since 1999 names AI EHR prompts: verify every AI-surfaced diagnosis
OIG's first MA compliance guidance since 1999 names AI EHR prompts and add-only chart reviews as risk-adjustment exposures. Verify every AI-surfaced dx.
Reviewed by Jess P., CPC
Published July 26, 2026

Key Takeaways
- →OIG's first Medicare Advantage-specific compliance guidance since 1999 (February 3, 2026) names using EHR queries "including prompts generated by artificial intelligence algorithms" to add unsupported risk-adjusting diagnoses as fraud and abuse conduct.
- →The same guidance treats add-only chart review as an exposure: failing to remove a previously submitted code that a chart review shows is unsupported is the violation OIG describes, not just a missed best practice.
- →In a simulated-encounter study across ambient AI scribe platforms, omissions were the most frequent error (about 76%), and the notes also produced meaning-reversal errors that flip what happened in the room.
- →Don't let "the study said scribes are fine" into your QA logic: the widely-cited NEJM AI trial (238 physicians) measured documentation time and burnout, never accuracy. Scribe error rates come from separate evaluation research.
- →Whether a coder, a chart review, or an algorithm surfaced a diagnosis, it needs its own MEAT support at the date of service before the HCC is submitted, or it becomes an overpayment.
A diagnosis can reach your queue two new ways that didn't exist a few years ago. An ambient AI scribe drafts it into the note from the visit audio, or an EHR prompt suggests it at the point of care. Neither one is documentation. On February 3, 2026, the HHS Office of Inspector General made that distinction a compliance problem.
OIG published its Medicare Advantage Industry Segment-Specific Compliance Program Guidance that day, the first compliance guidance the office has written specifically for MA since 1999. Its risk-adjustment section reads like a list of the shortcuts coders already push back on, and now it has AI in it by name.
What OIG actually named
The risk-adjustment section (part C, around page 20) lists conduct OIG treats as fraud and abuse. Four items land on the coder's desk, in OIG's own words:
- "querying physicians via electronic medical record platforms (including prompts generated by artificial intelligence algorithms) ... to add risk-adjusting diagnoses that patients did not have"
- "using chart reviews to identify additional diagnoses that increased risk scores inappropriately"
- "failing to remove diagnosis codes previously submitted to CMS when chart reviews provide information that those codes were unsupported or otherwise invalid"
- "providers submitting diagnoses that were not supported by the enrollees' medical records"
Read together, they point one way. OIG treats add-only chart review as an exposure, and it treats removing an unsupported code as a requirement, not a courtesy. OIG separately flags AI in utilization management, where an algorithm "may not be compliant ... if it determines coverage based on a larger data set instead of the individual patient's medical history." That is the prior-auth angle, adjacent to your desk, not the coding one.
The error you won't catch is the omission
The EHR prompt suggests a diagnosis. The ambient scribe writes the note the diagnosis has to be supported in. The failure mode that should change how you read a scribe note is the one you can't see: omission. In a simulated-encounter study across ambient platforms, omissions were the most common error, about 76% of all errors found. A condition gets addressed in the room, the scribe leaves it off the page, and nothing on the page looks wrong. The same research documents meaning-reversal errors, where the note flips what happened: one platform logged a patient as "alert" in a case involving altered mental status, another dropped a medication allergy the patient had reported. Reported overall error rates for modern ambient scribes run around 1% to 3%, small in the abstract and consequential in a chart.
Keep that error research separate from the efficiency research, because they answer different questions. A randomized trial in NEJM AI (Lukac et al., 238 outpatient physicians across 14 specialties) measured documentation time and burnout, not accuracy. It found Nabla cut time-in-note about 9.5% while Microsoft DAX Copilot showed no significant change. The time savings are small and product-specific. That trial says nothing about error rates, and you shouldn't cite it for any.
Two AI inputs, one check
Both inputs do the same thing to your chart: they add. Here is each one mapped to the check it forces.
| Where AI adds to the chart | What it puts there | Its main failure mode | Your verification step |
|---|---|---|---|
| Ambient AI scribe (drafts the note) | A full encounter note generated from the visit audio | Omission (a condition addressed but never written up); meaning-reversal (the note flips what happened) | Read for what's missing, then cross-check the assessment against orders, labs, and meds. Confirm anything the note negates or reverses against the rest of the record. |
| AI EHR prompt (suggests the dx) | A risk-adjusting diagnosis, surfaced at the point of care | Adds a diagnosis the record may not support | Treat it as suspecting, not a capture. Require provider documentation that addresses it at the DOS before it is codable. |
| Add-only chart review | Extra diagnoses pulled from the chart to raise the score | No removal of codes later shown unsupported | Confirm independent MEAT for each add. Flag any prior submission the review shows was unsupported for deletion. |
First, does it even map to an HCC?
Before you spend a minute chasing documentation, confirm the AI surfaced something that risk-adjusts at all. A prompt or a scribe will float diagnoses that carry no HCC in the payment-year model, and those never move a risk score no matter how cleanly they are documented. Run the code through the ICD-10-to-HCC crosswalk and check it against the V28 model before you treat it as a capture worth defending. If it doesn't map, it isn't a risk-adjustment question, and your verification time belongs on the codes that are.
The guidance cuts against "more codes"
Every one of those inputs pushes more diagnoses onto the chart. OIG's rule cuts the other way. An AI-prompted HCC with no record support is precisely what the guidance says has to come off, and failing to remove it is the violation OIG describes. OIG doesn't ban the tools. The unsupported code is the problem, whether a coder, a chart review, or an algorithm surfaced it.
It's the same check you already run. The only new thing is where the diagnosis came from, so every AI-surfaced diagnosis still needs its own MEAT before you submit the HCC. Walk it through the MEAT criteria and confirm the supporting evidence is in the record at the date of service. If it isn't, the code doesn't move the RAF, because it shouldn't be on the claim at all.
What coders should do now
- 1Treat any diagnosis that first appears because an EHR prompt suggested it, or a computer-assisted-coding tool auto-suggested the code, as suspecting, not captured; it is codable only once provider documentation addresses it at the date of service.
- 2Read AI-drafted notes for what is missing, not just what is wrong; for encounters where labs, meds, or orders show a condition was addressed, confirm the scribe actually documented it.
- 3Confirm anything an AI-drafted note negates or reverses ("denies," "no history of," "alert") against the rest of the chart before you let it drop or downcode a real condition.
- 4Run the remove side, not only the add side: when a chart review shows a previously submitted code was unsupported, flag it for deletion, because under this guidance leaving it on is the exposure.
- 5Line up the note, labs, meds, and orders for each AI-surfaced diagnosis in the [evidence builder](/evidence) so the MEAT (or the gap) is obvious before you submit the HCC.
Frequently Asked Questions
Does the OIG guidance ban AI scribes or AI EHR prompts?
No. It doesn't prohibit the tools. It names a specific misuse, using EHR queries "including prompts generated by artificial intelligence algorithms" to add diagnoses patients didn't have, as fraud and abuse conduct, and it makes removing unsupported codes an obligation. The tool is fine. The unsupported code isn't.
An EHR prompt suggested an HCC. Can I code it from the prompt?
Not from the prompt alone. Treat it the way you would treat a lab that suspects a condition: a lead, not evidence. You still need provider documentation that addresses the diagnosis at the date of service. Suspecting isn't coding.
The AI scribe drafted the note. Isn't a signed note still documentation?
Once the provider signs it, the note is the provider's documentation, which is exactly why the omission and meaning-reversal errors matter. If the signed note missed a condition or flipped one, the record doesn't support the code no matter how it was drafted. Verify against the rest of the chart.
Did the NEJM study find AI scribes cause coding errors?
No, and this one is worth keeping straight. The NEJM AI randomized trial (238 physicians) measured documentation time and burnout, not accuracy, and found the time savings were modest and product-dependent. The omission and meaning-reversal findings come from a separate body of evaluation research, not that trial.
Sources
- Medicare Advantage Industry Segment-Specific Compliance Program Guidance — HHS Office of Inspector General, Feb 3, 2026
- Ambient AI Scribes in Clinical Practice: A Randomized Trial — NEJM AI, Nov 26, 2025
- Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters — Mayo Clinic Proceedings: Digital Health, Oct 9, 2025
- Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice — npj Digital Medicine, Sep 24, 2025
Related Tools
Jess P., CPC
Certified Professional Coder
Jess reviews HCC Buddy editorial content for accuracy against the current CMS-HCC model and the active FY ICD-10-CM tabular release.
Get CMS Updates in Your Inbox
RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.
More from CMS Watch
Coding errors now drive nearly 7 in 10 completed payer-audit denials, MDaudit reports
July 26, 2026
PolicyAn ICD coding model scored 0.709 on a 50-code subset and 0.527 across all 15,761 codes
July 22, 2026
PolicyUnitedHealthcare commercial adds five lab policies September 1, and four cap how often it will pay
July 21, 2026

