Skip to content
Back to CMS Watch
PolicyJuly 21, 2026·6 min read

AI summaries often drop a note's diagnostic hedges. Outpatient rules make that a coding problem.

A new preprint reports that three LLMs preserved a clinical note's uncertainty cues poorly, often less than half the time, and that when they distort it they almost always collapse a hedge into a flat assertion. Section IV.H is what turns that into a code the guidelines say not to report.

AI documentationICD-10-CM Guidelinesrisk adjustmentdocumentation integrityoutpatient coding
Jess P., CPC

Reviewed by Jess P., CPC

Published July 21, 2026

A dictation microphone beside a closed patient folder and a red pen, illustrating diagnostic uncertainty lost when AI rewrites a note
The word that decides whether a diagnosis is codable outpatient is the one a summarizer is most likely to drop.Image: HCC Buddy

Key Takeaways

  • A benchmark preprint posted June 16, 2026 (arXiv:2606.18471) built 1,200 clinical documents with 9,184 uncertainty annotations across five levels and reported that the three general-purpose LLMs tested "preserve the original uncertainty cues poorly, often less than half the time" when summarizing or revising clinical text.
  • The authors report that when the models distort uncertainty they "almost always collapse hedged language into definite assertions," with partial shifts and over-hedging together accounting for less than 8% of retained pairs under the baseline prompt condition.
  • ICD-10-CM Official Guidelines Section IV.H bars coding diagnoses documented as probable, suspected, questionable, rule out, compatible with, consistent with, or working diagnosis in the outpatient setting, directing the coder to code signs, symptoms, or abnormal findings instead.
  • Section II.H sets the opposite rule for inpatient admissions to short-term acute, long-term care, and psychiatric hospitals, where an uncertain diagnosis documented at the time of discharge is coded as if established.
  • Because Sections II.H and IV.H run opposite rules, a collapsed hedge is neutral inpatient only where it lands in the discharge diagnostic statement, while outpatient it produces a diagnosis the documentation does not support at the required certainty.
  • The benchmark is built from MIMIC-IV-Note discharge summaries and radiology reports plus TCGA pathology reports, the word "outpatient" does not appear in the paper, and the 9,184 annotations came from a rule-based pipeline validated at 82% accuracy (Cohen's kappa 0.61) rather than full manual annotation.

A tool that turns "possible pneumonia" into "pneumonia" has made your coding decision for you, and in the outpatient world it made it wrong. A benchmark preprint posted June 16, 2026 put a number on how often that happens.

The paper is "Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text" (arXiv:2606.18471). The authors built a benchmark of 1,200 clinical documents carrying 9,184 uncertainty annotations across five levels, then ran three general-purpose LLMs (gpt-oss-120b, gemini-2.5-flash, and claude-haiku-4.5) at two jobs: summarizing clinical text, and revising it into plain language for patients. In the authors' words, the models "preserve the original uncertainty cues poorly, often less than half the time," and they "struggle with nuanced distinctions between adjacent levels."

The paper never mentions ICD-10-CM, diagnosis code assignment, or reimbursement. The coding consequence below is ours, not theirs.

Why a hedge word is a coding decision

The ICD-10-CM Official Guidelines for Coding and Reporting draw a hard line at exactly the words these models are dropping.

Section IV.H, the outpatient rule, reads: "Do not code diagnoses documented as 'probable', 'suspected,' 'questionable,' 'rule out,' 'compatible with,' 'consistent with,' or 'working diagnosis' or other similar terms indicating uncertainty. Rather, code the condition(s) to the highest degree of certainty for that encounter/visit, such as symptoms, signs, abnormal test results, or other reason for the visit."

So a note that says "possible pneumonia" gets the cough or the abnormal chest film. A note that says "pneumonia" gets pneumonia. One word decides it, and a summarizer that treats that word as filler has settled the question before a coder sees the chart.

The distortion runs in the dangerous direction

This is the finding that should worry an outpatient coder most, and it's the one the abstract doesn't advertise.

Under the baseline prompt condition, the authors report that partial certainty shifts and over-hedging together account for less than 8% of retained pairs, and conclude that when these models distort uncertainty "they almost always collapse hedged language into definite assertions rather than make subtle shifts along the uncertainty scale." The dominant error isn't a diagnosis going fuzzy. It's a hedge going flat.

Flat is the direction that manufactures a codable-looking diagnosis out of documentation that never supported one.

The authors also tested a "Guarded" prompt condition that instructs the model to rely only on source information and preserve hedged language. Their conclusion on it is worth quoting to a vendor verbatim: "Guarded prompting reduces but does not eliminate this distortion."

Inpatient runs the opposite rule

Section II.H says the reverse: "If the diagnosis documented at the time of discharge is qualified as 'probable,' 'suspected,' 'likely,' 'questionable,' 'possible,' or 'still to be ruled out,' 'compatible with,' 'consistent with,' or other similar terms indicating uncertainty, code the condition as if it existed or was established." Its own note limits it to "inpatient admissions to short-term, acute, long-term care and psychiatric hospitals."

Run the models' failure mode through both rules and they don't come out the same.

SettingGoverning guidelineModel collapses the hedge ("possible pneumonia" becomes "pneumonia")
Outpatient, physician officeIV.H, uncertain diagnoses are not codedA diagnosis gets coded that the guidelines say may not be coded from that documentation
Inpatient acute, LTCH, psychiatricII.H, an uncertain diagnosis at discharge is coded as establishedNeutral only where the change lands in the provider's own discharge diagnostic statement

That inpatient backstop is narrower than it looks, and three things fall outside it. II.H covers "the diagnosis documented at the time of discharge," so a hedge hardened in an H&P, a consult, a progress note, or the brief hospital course gets nothing from it. II.H also licenses coding the provider's own uncertain statement. A tool that drafts the discharge summary itself and writes "pneumonia" produces a diagnosis the provider never independently made, and if that draft is signed unread, the coder has no way to see it from the chart. And II.H's note names short-term acute, long-term care, and psychiatric hospitals, which leaves inpatient rehab, SNF, home health, and hospice out.

So inpatient has a partial backstop for one document. Outpatient has none. The physician-office and hospital-outpatient encounters that Medicare Advantage risk adjustment draws on run on the outpatient rule, which puts risk-adjustment coders on the exposed side.

The hedge lists in the two guidelines are not identical

Read the two sections side by side and something small shows up that matters when you build a screening list.

TermNamed in IV.H (outpatient)Named in II.H (inpatient)
probableyesyes
suspectedyesyes
questionableyesyes
rule out / still to be ruled out"rule out""still to be ruled out"
compatible withyesyes
consistent withyesyes
working diagnosisyesno
possibleno, falls under "other similar terms"yes
likelyno, falls under "other similar terms"yes

Both sections close with "other similar terms indicating uncertainty," so nothing here is a loophole. But "possible," the exact word the paper puts in its title, isn't spelled out in the outpatient section. If someone on your team builds a hedge-word screen by copying the outpatient list literally, one of the most common hedges in clinical writing won't be on it. Build the screen from the union of both lists.

What the paper tested, and what it did not

Keep the limits straight, because this is a preprint and the gap between what it measured and what your desk does is wide.

The benchmark is built from MIMIC-IV-Note discharge summaries and radiology reports (the assessment, brief hospital course, discharge diagnosis, impressions and findings sections) plus TCGA pathology reports. The word "outpatient" appears zero times in the paper. Nothing in the benchmark is a physician-office note, so the outpatient reading above is an inference from the guidelines rather than a tested result.

The 9,184 annotations also weren't hand-built. The authors used a rule-based extraction pipeline deliberately, to avoid having an LLM both generate and be graded on the labels, and validated a sample of it at 82% accuracy with an inter-reviewer agreement of Cohen's kappa 0.61, which they describe as moderate. They note some extraction errors may remain.

And the three models are off-the-shelf general-purpose LLMs. The paper does not test a named ambient scribe or any other clinical documentation product.

The certainty diff

Here's the check, since the paper stops at naming the problem.

Pull a sample of encounters where a summary, a revision, or a carried-forward assessment came out of a tool rather than the provider's own keystrokes. For each diagnosis in the output, find the same diagnosis in the signed source note and compare one thing: the certainty language.

Sort what you find into three piles.

  • Match. The output says what the provider said. Code it normally.
  • Hedge collapsed. The source hedges, the output is flat. Outpatient, this is the dangerous pile, and the paper says it's the pile to expect. The code comes off, and you fall back to the documented sign, symptom, or abnormal finding per IV.H.
  • Certainty softened. The source is definite, the output hedges. Outpatient, you've lost a supported diagnosis, so verify against the source note and code the source.

Anything you can't resolve from the record is a provider query, not a judgment call. Note which pile is bigger, too. A shop where collapsed hedges dominate has a compliance problem; the reverse is a revenue problem, and the two get fixed differently.

This is the same discipline as the capture-delta audit run against a different axis. That one asks which codes the tool added. This one asks whether the codes it kept still mean what the provider meant. Related coverage: an AI hit 91% predicting ICD categories, and its scoring counted sepsis for a UTI as correct, and the QA that still applies to AI-generated notes before you keep the HCC.

Where a human coder is still required

A certainty level isn't a data field. It lives in the provider's word choice, and the only way to know whether a summary preserved it is for a person to read the summary against the signed source note. The authors' own framing is that this is "a failure mode not captured by standard evaluation metrics," so no fluency score and no confidence number on the tool's output will flag it.

That read is also the read that decides whether a diagnosis is supported at all and whether the documentation behind it holds up when somebody asks why the code is on the claim. A tool that writes "pneumonia" where the provider wrote "possible pneumonia" hasn't made an error the tool can see. It's made one only a coder can see, and it's the coder's name on the claim.

What coders should do now

  1. 1Run a certainty diff on a live sample. For every diagnosis in an AI-summarized or AI-revised note, find the same diagnosis in the signed source note and compare only the certainty language. Log how often it matches, how often a hedge collapsed, and how often certainty softened.
  2. 2Expect the collapsed-hedge pile to be the big one. The authors report that is the dominant distortion, and outpatient it is the direction that puts an unsupported diagnosis on a claim.
  3. 3Build your hedge-word screen from the union of Sections II.H and IV.H. "Possible" and "likely" are spelled out in II.H but reach IV.H only through its "other similar terms indicating uncertainty" clause.
  4. 4Treat a collapsed hedge as an unsupported code outpatient. Pull the code, fall back to the documented sign, symptom, or abnormal finding, and check the encounter against the [MEAT criteria](/meat-criteria) before anything goes on the claim.
  5. 5Ask your documentation vendor one specific question: what does the tool do to preserve hedged language, and how was that measured. The paper's own guarded-prompting result is that instructing a model to preserve uncertainty "reduces but does not eliminate" the distortion, so "we prompt for it" is not an answer.

Frequently Asked Questions

Can you code a diagnosis documented as "possible" in an outpatient visit?

No. ICD-10-CM Official Guidelines Section IV.H directs that diagnoses documented as probable, suspected, questionable, rule out, compatible with, consistent with, working diagnosis, or other similar terms indicating uncertainty are not coded in the outpatient setting. Code the condition to the highest degree of certainty for that encounter instead, such as the sign, symptom, abnormal test result, or other reason for the visit.

Why can inpatient coders report an uncertain diagnosis when outpatient coders cannot?

Section II.H of the ICD-10-CM Official Guidelines says that when a diagnosis is qualified as uncertain at the time of discharge, the condition is coded as if it existed or was established. The guideline's own note limits it to inpatient admissions to short-term acute, long-term care, and psychiatric hospitals. Section IV.H then states the outpatient rule differs, and adds a note that this differs from the coding practices used by those hospitals.

Does AI note summarization change how certain a diagnosis looks?

A preprint benchmark (arXiv:2606.18471, posted June 16, 2026) reported that three general-purpose LLMs preserved the original uncertainty cues poorly, often less than half the time, when summarizing or revising clinical text, and that when they distorted uncertainty they almost always collapsed hedged language into definite assertions. Note the limits: it is a preprint, it tested off-the-shelf models rather than a clinical documentation product, and its documents are hospital discharge summaries, radiology reports, and pathology reports rather than office notes.

Does this apply to Medicare Advantage risk adjustment?

The physician-office and hospital-outpatient encounters that risk adjustment draws on follow the outpatient guideline, so Section IV.H applies and an uncertain diagnosis is not coded. Inpatient hospital encounters are also an acceptable risk-adjustment source and follow Section II.H, so the answer depends on the encounter type rather than on risk adjustment as a whole.

How do you check whether an AI tool is changing diagnostic certainty?

Compare the tool's output against the signed source note diagnosis by diagnosis, looking only at the certainty language, and sort the results into matches, collapsed hedges, and softened certainty. Fluency and coherence scores will not surface it. The authors describe this as a failure mode not captured by standard evaluation metrics, so the check has to be a human read of both documents side by side.

Related topics:AI documentationICD-10-CM Guidelinesrisk adjustmentdocumentation integrityoutpatient coding
Jess P., CPC

Jess P., CPC

Certified Professional Coder

Jess reviews HCC Buddy editorial content for accuracy against the current CMS-HCC model and the active FY ICD-10-CM tabular release.

Get CMS Updates in Your Inbox

RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.