Skip to content
Back to CMS Watch
PolicySeptember 7, 2026·5 min read

In a coding study, AI hit 91.5% on CPT but 23.9% on ICD-10, with laterality the top error

Public large language models coded 90 hand-surgery charts at 91.5% correct on CPT but only 23.9% on ICD-10, and the authors report the single most common ICD-10 error was dropped or wrong laterality. Left, right, or bilateral is a specificity call the Official Guidelines still put on the coder.

AI codingICD-10-CMlateralitydocumentationcoding accuracy
HCC Buddy

By the HCC Buddy Coding Team

Published September 7, 2026

Left and right hand X-rays on a radiology lightbox beside a blank coding worksheet, showing the ICD-10 laterality AI coding missed.
Illustration: HCC Buddy

Key Takeaways

  • In a study published in the Journal of the American Academy of Orthopaedic Surgeons (August 1, 2026; epub May 22, 2026), public large language models coded 90 deidentified hand-surgery charts at 91.5% correct on CPT procedure codes but only 23.9% correct on ICD-10 diagnosis codes.
  • The authors report the single most common ICD-10 error was incorrect or omitted laterality, and that prompts rewritten to emphasize laterality improved accuracy (40%).
  • The study reports no meaningful difference in accuracy by note author or by prompt style (zero-shot, one-shot, multishot, or chain-of-thought), and concluded the models are not ready for independent coding use.
  • The FY2026 ICD-10-CM Official Guidelines (Section I.B.13) state that unspecified-side codes should rarely be used, and allow the coder to assign laterality from another clinician's documentation or query the provider when the side is unclear.
  • The three procedures studied included carpal tunnel and cubital tunnel release, whose ICD-10 codes (for example carpal tunnel G56.00 unspecified versus G56.01 right, G56.02 left, and G56.03 bilateral) differ only by laterality.

On August 1, 2026, the Journal of the American Academy of Orthopaedic Surgeons published a study that asked a narrow question: can a public large language model read a clinical note and assign the correct billing codes? The models handled CPT procedure codes well and ICD-10 diagnosis codes poorly, and the authors report the single most common ICD-10 error was one thing a coder checks by reflex. Laterality. Left, right, or bilateral.

What the study tested

The researchers pulled 90 deidentified charts, evenly split across three orthopaedic hand surgeons and three procedures: cubital tunnel release, carpal tunnel release, and trigger finger release. They recorded the correct ICD-10 diagnosis code and CPT procedure code for each, then asked three models (ChatGPT 3.5, ChatGPT 4.0, and Gemini) to code the notes. Each note was posed four ways: zero-shot, one-shot, multishot, and chain-of-thought prompting. The study reports the work was limited to hand-surgery clinic and operative notes.

CPT held up, ICD-10 did not

Across all three models, CPT came back 91.5% correct and ICD-10 only 23.9%, the study reports. Two other results are worth a coder's attention. Neither the note's author nor the prompt style moved the numbers, and ChatGPT 3.5 produced less accurate ICD-10 codes than ChatGPT 4.0 or Gemini. The authors' conclusion is blunt: public-facing models "require additional optimization to interpret clinical documentation for coding purposes and are not ready for independent use."

What the models codedCorrect, as the study reports it
CPT procedure codes91.5%
ICD-10 diagnosis codes23.9%
ICD-10 after laterality-focused promptingimproved accuracy (40%)

Laterality was the sticking point

The paper names the failure mode directly: the most common ICD-10 error was incorrect or omitted laterality. When the authors rewrote the prompt to emphasize laterality specifically, they report accuracy improved (40%). Read that against what the FY2026 ICD-10-CM Official Guidelines actually require. Section I.B.13 says unspecified-side codes "should rarely be used," reserved for records where the side is not documented and cannot be clarified. A model that returns the unspecified-side code when the note says "right" is not being imprecise. It is producing a code the Guidelines tell you to avoid.

The codes behind the finding

Two of the three procedures the study coded live in the G56 family, where laterality is the only thing separating a specific code from an unspecified one. This is the shape of the error a coder is checking for, and it's the same category-versus-code gap a separate study hit when a model predicted the diagnosis bucket but not the billable code.

ICD-10 codeDescriptorSpecificity
G56.00Carpal tunnel syndrome, unspecified upper limbunspecified side
G56.01Carpal tunnel syndrome, right upper limbspecific
G56.02Carpal tunnel syndrome, left upper limbspecific
G56.03Carpal tunnel syndrome, bilateral upper limbsspecific
G56.20Lesion of ulnar nerve, unspecified upper limbunspecified side
G56.21Lesion of ulnar nerve, right upper limbspecific

Cubital tunnel (lesion of ulnar nerve) mirrors carpal tunnel: G56.20 unspecified, G56.21 right, G56.22 left, G56.23 bilateral. The distance between the top row and the rest is exactly the distance an AI draft skips.

What the study did not test

This is 90 hand-surgery charts and three general-purpose chatbots, not a coding department. It didn't test the fine-tuned vendor coding tools or ambient scribes a practice actually buys. It didn't test the current model versions, and it says nothing about outpatient risk-adjustment charts or any other specialty. Take the 91.5% CPT result the same way: it graded three well-defined hand procedures, not the multi-code operative notes where CPT selection gets genuinely hard. The generalizable lesson isn't that AI is bad at coding. It's that these tools drop the specific detail, and laterality runs through much of the code set, from the injury and musculoskeletal chapters to the eye and ear.

Where the coder still owns the code

The study's own fix (tell the model to watch laterality) is a coder's habit, not a feature. When you review an AI-suggested or scribe-generated code, the laterality character is the first thing to check against the note, because it is the first thing the model gets wrong. If the provider did not document the side, the Guidelines let you code it from another clinician's documentation, and query the provider when the record conflicts, rather than defaulting to the unspecified code. That judgment (read the record, apply I.B.13, decide when to query) is the part no model in this study could do on its own.

What coders should do now

  1. 1On any AI-suggested or scribe-generated ICD-10 code, check the laterality character against the note before you accept it. The study reports this is where the models miss most, and it is a two-second check against the record.
  2. 2Reserve unspecified-side codes for records that genuinely do not state a side. The FY2026 Guidelines (I.B.13) say they should rarely be used, so an AI default to the unspecified code is an error to correct, not documentation you inherit.
  3. 3When the provider's note omits the side but another clinician documented it, code the side from that documentation as the Guidelines allow, and query the provider when the record conflicts.
  4. 4Look up the side-specific code directly in the [code book](/code-book) or [encoder](/encoder) instead of shipping the unspecified default an AI draft hands you.
  5. 5Do not read the study's 91.5% CPT result as a free pass. It graded three hand procedures, not your specialty or your multi-code operative notes, so re-check CPT the same way.

Frequently Asked Questions

Is AI worse at ICD-10 or CPT coding?

In this hand-surgery study, the models were far worse at ICD-10 (23.9% correct) than at CPT (91.5% correct). The result is specific to three procedures and three general-purpose models, so treat it as a signal about where AI coding is weak, not a universal benchmark.

What is the most common AI error in ICD-10 coding, and why does it matter?

In this study the single most common ICD-10 error was incorrect or omitted laterality: the model returned an unspecified-side code, or the wrong side, when the note supported a specific one. Under the FY2026 Guidelines an unspecified-side code should rarely be used, so an AI draft that defaults to it hands the coder a specificity gap to close, not a finished code.

When can I use an unspecified-laterality ICD-10 code?

Under the FY2026 ICD-10-CM Official Guidelines (Section I.B.13), unspecified-side codes should rarely be used, such as when the record does not document the side and it cannot be clarified. If another clinician documented the side you may code from that; if the record conflicts, query the provider.

Does better prompting fix AI laterality errors?

The study found prompt style (zero-shot, one-shot, multishot, chain-of-thought) made no meaningful difference, but a prompt rewritten to emphasize laterality improved accuracy (40%). Even then the authors concluded the models are not ready for independent coding use.

Does this study apply to risk-adjustment or outpatient coding?

No. It tested 90 hand-surgery clinic and operative notes with ChatGPT 3.5, ChatGPT 4.0, and Gemini. It did not test vendor coding tools, ambient scribes, current model versions, or outpatient and risk-adjustment charts.

Related topics:AI codingICD-10-CMlateralitydocumentationcoding accuracy
HCC Buddy

HCC Buddy Coding Team

Editorial

Every HCC Buddy news article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.

Get CMS Updates in Your Inbox

RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.