Skip to content
Back to CMS Watch
PolicyJuly 28, 2026·6 min read

General-purpose AI coded complex surgery at 5% to 12.5% exact-match, a new study reports

A peer-reviewed study reports out-of-the-box LLMs scored 5% to 32.5% exact-match on multi-code operative reports, while a fine-tuned vendor tool hit 80% and external auditors 22.5%. The study was authored by the tool's maker. Here is how these coding-accuracy numbers are built, and what a coder still checks.

AI documentationcomputer-assisted codingCPTsurgical codingcoding accuracy
HCC Buddy

By the HCC Buddy Coding Team

Published July 28, 2026

Operative-report folders and a magnifying glass on a desk beside a laptop, illustrating exact-match CPT coding accuracy testing of AI tools
A July 2026 study graded four coding methods on exact-match CPT code sets. The bar was every code, no omissions and no extras.Image: HCC Buddy

Key Takeaways

  • A study published in Plastic and Reconstructive Surgery - Global Open on July 15, 2026 graded CPT coding on 120 plastic and reconstructive operative reports, 40 each at low, medium, and high complexity, against an exact-match reference standard set by a four-member expert consensus panel.
  • The study reports that out-of-the-box general-purpose LLMs scored 5.0% to 12.5% exact-match on the high-complexity operative reports (four or more CPT codes), and 21.7% to 35.8% across all 120 cases combined.
  • The study reports that a team of external professional medical coders scored 42.5% exact-match overall and 22.5% on the high-complexity cases, under a rule that counted a code set correct only if every code matched with no omissions and no extra codes, with modifiers excluded from the analysis.
  • The fine-tuned vendor tool (ProCode) scored 86.7% overall and 80% on high-complexity cases, the highest of any method tested. Six of the study's eight authors are affiliated with ProCode, Inc., the company that makes the tool.
  • Even the highest-scoring method missed the exact code set on 8 of 40 high-complexity operative reports, and the study excluded modifiers and code order from the accuracy definition entirely.

On July 15, 2026, the journal Plastic and Reconstructive Surgery - Global Open published a comparative study grading four ways of assigning CPT codes to surgical operative reports: three off-the-shelf large language models, a team of external professional coders, and a fine-tuned vendor tool. Every method was scored against a reference standard built by a four-member expert consensus panel, and the bar was exact agreement on the full code set.

The number a vendor slide will pull from this is the fine-tuned tool's 86.7%. The numbers worth a coder's time are all the others, and the footnotes under all of them.

What the study measured, and how it graded

The researchers took 120 de-identified plastic and reconstructive operative reports and split them evenly into three complexity tiers: low (1 to 2 CPT codes), medium (3 codes), and high (4 or more codes), 40 reports each. Four methods coded every report: baseline general-purpose LLMs, an external professional medical auditing service, and a fine-tuned hybrid tool the authors call ProCode. A panel of two board-certified plastic surgeons and two senior billing specialists, each with more than 20 years of experience, set the expert consensus reference standard.

Accuracy was defined strictly. A code set counted as correct only if every CPT code matched the reference with no omissions and no additions. Code order and duplicates were ignored, modifiers were excluded from the analysis, and no partial credit was given. The authors chose that exact-match rule to mirror real billing, where an incomplete or padded code set gets a claim denied.

The full accuracy table

Here is every method against every complexity tier, as the study reports it. The percentages are exact-match on the full CPT code set.

MethodLow (n=40)Medium (n=40)High (n=40)Overall (n=120)
GPT-5.0 (OpenAI)62.5%32.5%12.5%35.8%
Gemini 2.5 Pro (Google)62.5%25.0%7.5%31.7%
Claude Sonnet 4.5 (Anthropic)47.5%12.5%5.0%21.7%
External auditors65.0%40.0%22.5%42.5%
ProCode (fine-tuned)100.0%80.0%80.0%86.7%

Two patterns jump out. Every method except the fine-tuned tool falls off a cliff as the code count climbs, and the general-purpose models fall hardest. On the high-complexity reports, the three off-the-shelf models the study tested landed between 5.0% and 12.5%.

Why the general-purpose models collapse on complex cases

The story in that first column is the one that matters most for a working coder. A single-procedure operative note is a fill-in-the-blank task, and a general model handles it more than half the time. A multi-code reconstruction is not. It needs the add-on logic, the bundling rules, and the specialty-specific exclusions that live in the CPT book and in a coder's head, not in the narrative text.

The study's own framing is blunt about this: the models were run out of the box, single pass, no human correction and no iterative prompting, because that is how a rushed office actually uses them. Under those conditions, a general chatbot got the full code set right on 1 in 20 to 1 in 8 complex operative reports. That is the number to keep when someone suggests pointing a general AI at your surgical queue.

Read the auditor number with the method next to it

The line that will get quoted out of context is the external auditors at 42.5% overall, and 22.5% on the hard cases. Taken alone it reads as "professional coders are worse than the AI." Read with the method, it says something narrower.

The auditors coded single pass, against a reference standard the study's own panel built, on a held-out set drawn from the vendor's data, and were graded on all-or-nothing full-set matching with no ability to query the surgeon. Real surgical coders send the operative report back with questions, work from the actual chart, and are judged code by code, not on a binary set match. Strip a coder of the query and grade the whole set as one pass-or-fail unit, and the number drops. The study measured a constrained task, not a coder's job. It is the same trap as grading an AI against the billing codes already on the claim: the yardstick decides the score.

The tool that won was built by the authors

The fine-tuned tool posted the best figure at every tier. It was also trained on roughly 55,000 operative reports, and, most important for how you read the result, six of the study's eight authors are affiliated with ProCode, Inc., the company that sells it. The tool being evaluated, the training data, and the paper making the claim all trace back to the same source. That does not make the 86.7% false. It means the number carries the weight of a vendor benchmark, not an independent one, and it belongs in the "the vendor reports" column until someone outside the company reproduces it.

Even taken at face value, the winning method missed the exact code set on 8 of 40 high-complexity reports, and modifiers, the part of surgical coding most likely to trigger a denial, were excluded from the scoring entirely.

Where a human coder is still required

Strip the branding off this study and the durable lesson is one every risk-adjustment and provider-office coder already lives: AI accuracy is a function of complexity, and the complex chart is exactly where the money and the audit risk sit. A tool that nails a single-procedure note can miss the multi-code reconstruction, the add-on service, the bundling exclusion, or the modifier that moves a claim.

So the coder's job does not disappear when an AI tool arrives on the desk. It moves to the code set as a whole. Did the tool drop a reportable code, or add one the note does not support? Are the add-ons and modifiers right, since this study did not even grade them? Does the documentation carry the procedure the code claims? Those are judgments about a record, made by someone who can be held to them, and no exact-match percentage in a vendor paper answers a single one. When one lands on your desk, ask what it was graded against and at what granularity, then look the codes up and read the note.

What coders should do now

  1. 1When a vendor quotes a coding-accuracy figure, ask four questions before anything else: exact full-set match or per-code, are modifiers included, whose reference standard was it graded against, and was it single-pass or with a provider query. This study's 42.5% auditor number and 86.7% tool number came from the same all-or-nothing rule with modifiers excluded.
  2. 2Check who built the benchmark. A tool that beats auditors on test data drawn from its own maker's pool, graded against a reference standard the same study built, has not been shown to beat them on your charts. Treat a vendor-authored accuracy claim as a vendor benchmark until an independent group reproduces it.
  3. 3Point your QA at the full code set, not individual codes. Both an omission and an extra code deny a claim, so review AI-suggested output for dropped reportable codes, unsupported added codes, add-on services, and modifiers, the last of which many accuracy studies (including this one) exclude.
  4. 4Treat complexity as the failure zone. General-purpose LLMs in this study scored above 60% on simple cases and 5% to 12.5% on the four-plus-code cases. Re-check every line on multi-code and add-on charts; that is where any AI tool degrades most.
  5. 5Apply the same read to risk-adjustment AI. An HCC tool that captures the single-condition encounter can still miss the multi-condition chart, and agreeing with a code is not the same as the record supporting it. Verify the encounter against the documentation before you trust a suggestion.

Frequently Asked Questions

Can AI code CPT better than a professional coder?

One July 2026 study reports that a fine-tuned vendor tool scored higher on exact-match CPT accuracy than a team of external auditors across all complexity tiers. But the auditors were graded single-pass on all-or-nothing full-set matching with modifiers excluded and no ability to query the surgeon, and the study was authored by the tool's maker. The comparison measured a constrained task, not a coder's full workflow.

What accuracy did ChatGPT and other general AI get on surgical CPT coding?

The study reports out-of-the-box general-purpose LLMs scored 47.5% to 62.5% exact-match on low-complexity operative reports (1 to 2 codes) and 5.0% to 12.5% on high-complexity reports (4 or more codes), run single-pass with no human correction. Overall across 120 cases they scored 21.7% to 35.8%.

Why did the professional auditors only score 42.5%?

Accuracy was defined as an exact match on the entire CPT code set, correct only if every code matched with no omissions and no extras, and no partial credit. The auditors coded single-pass against a reference standard the study's own panel built, without the provider queries real coders use. That all-or-nothing rule, not the coders' skill, is what produces a low percentage.

What does exact-match CPT accuracy mean?

In this study, a code set was counted correct only if all CPT codes matched the reference standard with no missing and no extra codes. Code order and duplicates were ignored and modifiers were excluded from the analysis. The authors chose exact-match because in real billing an incomplete or padded code set can trigger a denial or lost revenue.

Was the study independent of the vendor?

No. Six of the study's eight authors are affiliated with ProCode, Inc., the company that makes the fine-tuned tool that scored highest. The tool, its training data, and the paper trace back to the same source, so the 86.7% figure should be read as a vendor benchmark rather than an independent result until reproduced.

Related topics:AI documentationcomputer-assisted codingCPTsurgical codingcoding accuracy
HCC Buddy

HCC Buddy Coding Team

Editorial

Every HCC Buddy news article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.

Get CMS Updates in Your Inbox

RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.