What clinicians correct most in AI scribe drafts is orders and meds, a 2026 study finds
A JAMIA content analysis tagged 713 edits clinicians made to ambient AI note drafts. The most-edited content was orders, tests, and medications, not wording. For a provider-office coder, that's the order and drug narrative you reconcile against the record before a CPT or HCPCS line bills.
By the HCC Buddy Coding Team
Published September 9, 2026

Key Takeaways
- →A qualitative content analysis published in the Journal of the American Medical Informatics Association compared ambient AI-drafted note sections against the versions clinicians signed, tagging 713 edits across 314 draft and final note-section pairs from 200 clinical encounters, 33 specialties, 73 clinicians, and 2 vendor systems.
- →The most-edited content was clinical facts: procedures, tests, and lab orders were tagged in 40.0% of the 713 edits, symptoms in 30.3%, medications in 27.3%, and diagnoses in 25.9%. Because one edit could fall in more than one category, the shares total more than 100%.
- →Certainty and hedging edits were rare by comparison, 29 of the 713 tagged edits, or 4.1%, one of the least common categories and well below the clinical-content edits.
- →The study reports what clinicians chose to change, not whether the AI draft was right or wrong. It is a qualitative sample of 200 encounters on 2 vendor systems at one health system, so the frequencies describe that sample; it did not benchmark vendors or measure how many edits would have changed a billed code.
The AI scribe draft and the note your provider signs aren't the same document, and a JAMIA study published May 12, 2026 counts where they split: orders, tests, and medications, the exact lines a CPT or HCPCS claim is built on. Researchers at the University of California, Irvine tagged every change clinicians made across 314 draft and final note-section pairs, and the edits pile up on clinical facts, not wording.
What the study counted
The method sets how much weight the numbers carry. This is a qualitative content analysis of paired drafts and signed notes, not a vendor accuracy benchmark and not an audit of coded claims. Nobody scored the drafts right or wrong. Eight researchers built a framework of 11 edit categories across three groups, clinical content, terminology, and language style, then tagged every edit in the sample of 200 encounters. They logged 713 tagged edits in all.
One detail governs how you read every percentage below. Because a single edit could fall in more than one category, the study reports each figure as the share of the 713 tagged edits carrying that label, and the authors say plainly the shares don't sum to 100%. So a 40.0% isn't two in five note sections. It's 285 of the 713 tagged edits.
What clinicians fixed most
| Edit category | Count | Share of the 713 tagged edits |
|---|---|---|
| Procedures, tests, and lab orders | 285 | 40.0% |
| Symptoms | 216 | 30.3% |
| Medications | 195 | 27.3% |
| Diagnoses | 185 | 25.9% |
| Past medical history | 135 | 18.9% |
| Social context | 127 | 17.8% |
| Demographics and identifiers | 63 | 8.8% |
| Lay-to-jargon substitution | 51 | 7.2% |
| Abbreviation change | 32 | 4.5% |
| Certainty and hedging | 29 | 4.1% |
| Empathy and rapport | 22 | 3.1% |
The ranking is the finding. Orders, symptoms, medications, and diagnoses sit at the top. Wording, abbreviations, certainty, and tone sit at the bottom. The cosmetic edits are the rare ones. The ones that touch what a claim is built from are everywhere.
Why orders and meds are the coding exposure
Put the top of that table next to the coder's desk. Procedures, tests, and lab orders are the single most-edited content, and medications are close behind. Those are exactly the elements a professional or facility claim is built on, the CPT and HCPCS lines and the medication record. When a clinician swaps out what the draft said an encounter ordered, the draft narrative and the billable event have parted ways, and it's the draft narrative a downstream reader inherits.
That makes order and medication content the reconciliation you don't skip on an ambient-drafted encounter. A CPT or HCPCS code has to reflect the service actually performed and ordered, documented in the record, not a fluent sentence a model produced from the room audio. The check is the same one a careful biller already runs, now aimed at a draft: match the order and result content in the note against the order and result record itself, rather than trusting the narrative. In procedure-heavy specialties, where that content is densest, weight the review there first.
Diagnoses got edited too, but certainty barely moved
Diagnoses were edited in 25.9% of tagged edits, and revising the certainty of diagnostic and causal statements was one of five recurring edit themes the authors named. That side of the note matters for risk adjustment, and this desk has covered it. An over-confident draft that reads more certain than the visit was is a documented risk to how a diagnosis gets coded, and the mirror problem, a draft that quietly drops a chronic condition and costs a recapture, is the other half of it.
Keep it in proportion, though. Certainty and hedging edits were a small slice here, 29 of 713 tagged edits, 4.1%. That's consistent with the earlier finding that certainty drift is real but marginal in what clinicians touch. The bigger, less-discussed story in this study is the clinical-content edits above it, the orders, tests, and medications, which is what this piece is about.
Where a human coder is still required
The study stops at what clinicians edited, so treat what follows as our reading rather than its finding. The draft and the signed note both persist, so the gap between them is checkable before a claim goes out rather than after a payer asks. Matching the order and medication content against the actual order record is a review step, not a software feature. A tool that drafts from audio doesn't know which order was actually placed, and the clinician who corrected it was making a clinical call, not a coding one.
So the reviewer's job is the same as ever, just better aimed. The categories this study ranks tell you where an ambient draft drifts from the record, and that's where to spend the review. It's the discipline of confirming a condition was genuinely evaluated and addressed at the encounter, applied to a note a machine wrote first.
What this study does not establish
Be careful quoting these in a meeting. It's 200 encounters on 2 vendor systems at one health system, analyzed qualitatively, with moderate agreement among the reviewers who tagged the edits. The frequencies describe that sample. They aren't an industry error rate or a vendor comparison, and the paper never asks how many edits would have changed a billed code. What it does establish is narrower and more useful: clinicians revise specific categories of content at measurable rates, and the categories they revise most are the ones a code depends on.
That also puts a count under an argument this desk has made without one. The case that an AI-generated note still needs coder review before you keep what's on it used to rest on how the notes read. Now there's a tally of what the review step is catching.
What coders should do now
- 1On ambient-drafted encounters, reconcile the order, test, and medication content in the note against the actual order and result record before a CPT or HCPCS line bills. Procedures, tests, and lab orders were the most-edited content in the study, so that's where the draft narrative is least likely to match what was ordered.
- 2Weight your ambient-draft audit sample toward procedure-heavy specialties. That's where order and result content is densest, and where a mismatch between the draft and the record is most likely to reach a claim.
- 3Keep a lighter but real check on diagnostic certainty and dropped chronic conditions on the risk-adjustment side, and treat those as separate reviews from the order reconciliation. The certainty edit is a small share of what clinicians change, but it's the one that can flip an outpatient diagnosis from reportable to not.
- 4When the note's order or medication content doesn't match the record, query the provider to confirm what was actually ordered rather than coding the draft sentence as written.
Frequently Asked Questions
What did clinicians edit most in ambient AI note drafts?
In the JAMIA study, the most-edited content was procedures, tests, and lab orders, tagged in 40.0% of the 713 edits, followed by symptoms at 30.3%, medications at 27.3%, and diagnoses at 25.9%. The figures come from 314 draft and final note-section pairs across 200 encounters and describe that sample.
Why do order and medication edits matter for coding and billing?
A CPT or HCPCS code must reflect the service actually performed and ordered as documented in the record, and a medication line must match what was prescribed. When a clinician corrects the order or medication content between an ambient draft and the signed note, the draft narrative no longer matches the billable event, so reconciling that content against the order and result record before billing is the safeguard.
How often did clinicians change diagnostic certainty in the drafts?
Certainty and hedging edits were a small category: 29 of the 713 tagged edits, or 4.1%. Revising the certainty of diagnostic and causal statements was one of five recurring edit themes the authors named, but by frequency it was one of the least common categories, well below the clinical-content edits like orders, symptoms, and medications.
Does this study say ambient AI scribes are inaccurate?
No. The researchers ran a qualitative content analysis of what clinicians chose to edit between the draft and the signed note. They didn't score drafts as correct or incorrect, didn't benchmark the two vendor systems against each other, and didn't report how many edits would have changed a billed code.
Can I cite these percentages as an industry error rate?
No. The study covers 200 encounters, 314 note-section pairs, 33 specialties, 73 clinicians, and 2 vendor systems at one health system, analyzed qualitatively. The frequencies are the share of 713 tagged edits in that sample. Presenting them as a general accuracy or error rate overstates what the paper supports.
Sources
- What do clinicians edit in ambient AI-drafted clinical documentation? A qualitative content analysis — Journal of the American Medical Informatics Association (JAMIA), via PubMed Central, May 12, 2026
- Clinicians' rationale for editing ambient AI-drafted clinical notes: persistent challenges and implications for improvement — Journal of the American Medical Informatics Association (JAMIA), via PubMed Central, Jul 1, 2026
- ICD-10-CM Official Guidelines for Coding and Reporting FY 2026, Section IV.H (Uncertain diagnosis, outpatient) — CMS, Oct 1, 2025
Related Tools
Evidence Check
Work through what a note actually establishes before you keep an order, a drug, or a diagnosis off it.
MEAT Criteria
Confirm the condition was monitored, evaluated, assessed, or treated at the encounter you are coding.
Code Book
Check the official guideline text yourself rather than relying on a summary of it.
HCC Buddy Coding Team
Editorial
Every HCC Buddy news article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.
Get CMS Updates in Your Inbox
RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.
More from CMS Watch
AI scribe drafts can read more certain than the visit was, a 2026 study finds
September 8, 2026
PolicyIn a coding study, AI hit 91.5% on CPT but 23.9% on ICD-10, with laterality the top error
September 7, 2026
PolicyThe 2026 ACDIS/AHIMA query standard is final: the outpatient and HCC query rules it sets
September 5, 2026

