In a coding study, AI hit 91.5% on CPT but 23.9% on ICD-10, with laterality the top error
Public large language models coded 90 hand-surgery charts at 91.5% correct on CPT but only 23.9% on ICD-10, and the authors report the single most common ICD-10 error was dropped or wrong laterality. Left, right, or bilateral is a specificity call the Official Guidelines still put on the coder.
By the HCC Buddy Coding Team
Published September 7, 2026

Key Takeaways
- →In a study published in the Journal of the American Academy of Orthopaedic Surgeons (August 1, 2026; epub May 22, 2026), public large language models coded 90 deidentified hand-surgery charts at 91.5% correct on CPT procedure codes but only 23.9% correct on ICD-10 diagnosis codes.
- →The authors report the single most common ICD-10 error was incorrect or omitted laterality, and that prompts rewritten to emphasize laterality improved accuracy (40%).
- →The study reports no meaningful difference in accuracy by note author or by prompt style (zero-shot, one-shot, multishot, or chain-of-thought), and concluded the models are not ready for independent coding use.
- →The FY2026 ICD-10-CM Official Guidelines (Section I.B.13) state that unspecified-side codes should rarely be used, and allow the coder to assign laterality from another clinician's documentation or query the provider when the side is unclear.
- →The three procedures studied included carpal tunnel and cubital tunnel release, whose ICD-10 codes (for example carpal tunnel G56.00 unspecified versus G56.01 right, G56.02 left, and G56.03 bilateral) differ only by laterality.
On August 1, 2026, the Journal of the American Academy of Orthopaedic Surgeons published a study that asked a narrow question: can a public large language model read a clinical note and assign the correct billing codes? The models handled CPT procedure codes well and ICD-10 diagnosis codes poorly, and the authors report the single most common ICD-10 error was one thing a coder checks by reflex. Laterality. Left, right, or bilateral.
What the study tested
The researchers pulled 90 deidentified charts, evenly split across three orthopaedic hand surgeons and three procedures: cubital tunnel release, carpal tunnel release, and trigger finger release. They recorded the correct ICD-10 diagnosis code and CPT procedure code for each, then asked three models (ChatGPT 3.5, ChatGPT 4.0, and Gemini) to code the notes. Each note was posed four ways: zero-shot, one-shot, multishot, and chain-of-thought prompting. The study reports the work was limited to hand-surgery clinic and operative notes.
CPT held up, ICD-10 did not
Across all three models, CPT came back 91.5% correct and ICD-10 only 23.9%, the study reports. Two other results are worth a coder's attention. Neither the note's author nor the prompt style moved the numbers, and ChatGPT 3.5 produced less accurate ICD-10 codes than ChatGPT 4.0 or Gemini. The authors' conclusion is blunt: public-facing models "require additional optimization to interpret clinical documentation for coding purposes and are not ready for independent use."
| What the models coded | Correct, as the study reports it |
|---|---|
| CPT procedure codes | 91.5% |
| ICD-10 diagnosis codes | 23.9% |
| ICD-10 after laterality-focused prompting | improved accuracy (40%) |
Laterality was the sticking point
The paper names the failure mode directly: the most common ICD-10 error was incorrect or omitted laterality. When the authors rewrote the prompt to emphasize laterality specifically, they report accuracy improved (40%). Read that against what the FY2026 ICD-10-CM Official Guidelines actually require. Section I.B.13 says unspecified-side codes "should rarely be used," reserved for records where the side is not documented and cannot be clarified. A model that returns the unspecified-side code when the note says "right" is not being imprecise. It is producing a code the Guidelines tell you to avoid.
The codes behind the finding
Two of the three procedures the study coded live in the G56 family, where laterality is the only thing separating a specific code from an unspecified one. This is the shape of the error a coder is checking for, and it's the same category-versus-code gap a separate study hit when a model predicted the diagnosis bucket but not the billable code.
| ICD-10 code | Descriptor | Specificity |
|---|---|---|
| G56.00 | Carpal tunnel syndrome, unspecified upper limb | unspecified side |
| G56.01 | Carpal tunnel syndrome, right upper limb | specific |
| G56.02 | Carpal tunnel syndrome, left upper limb | specific |
| G56.03 | Carpal tunnel syndrome, bilateral upper limbs | specific |
| G56.20 | Lesion of ulnar nerve, unspecified upper limb | unspecified side |
| G56.21 | Lesion of ulnar nerve, right upper limb | specific |
Cubital tunnel (lesion of ulnar nerve) mirrors carpal tunnel: G56.20 unspecified, G56.21 right, G56.22 left, G56.23 bilateral. The distance between the top row and the rest is exactly the distance an AI draft skips.
What the study did not test
This is 90 hand-surgery charts and three general-purpose chatbots, not a coding department. It didn't test the fine-tuned vendor coding tools or ambient scribes a practice actually buys. It didn't test the current model versions, and it says nothing about outpatient risk-adjustment charts or any other specialty. Take the 91.5% CPT result the same way: it graded three well-defined hand procedures, not the multi-code operative notes where CPT selection gets genuinely hard. The generalizable lesson isn't that AI is bad at coding. It's that these tools drop the specific detail, and laterality runs through much of the code set, from the injury and musculoskeletal chapters to the eye and ear.
Where the coder still owns the code
The study's own fix (tell the model to watch laterality) is a coder's habit, not a feature. When you review an AI-suggested or scribe-generated code, the laterality character is the first thing to check against the note, because it is the first thing the model gets wrong. If the provider did not document the side, the Guidelines let you code it from another clinician's documentation, and query the provider when the record conflicts, rather than defaulting to the unspecified code. That judgment (read the record, apply I.B.13, decide when to query) is the part no model in this study could do on its own.
What coders should do now
- 1On any AI-suggested or scribe-generated ICD-10 code, check the laterality character against the note before you accept it. The study reports this is where the models miss most, and it is a two-second check against the record.
- 2Reserve unspecified-side codes for records that genuinely do not state a side. The FY2026 Guidelines (I.B.13) say they should rarely be used, so an AI default to the unspecified code is an error to correct, not documentation you inherit.
- 3When the provider's note omits the side but another clinician documented it, code the side from that documentation as the Guidelines allow, and query the provider when the record conflicts.
- 4Look up the side-specific code directly in the [code book](/code-book) or [encoder](/encoder) instead of shipping the unspecified default an AI draft hands you.
- 5Do not read the study's 91.5% CPT result as a free pass. It graded three hand procedures, not your specialty or your multi-code operative notes, so re-check CPT the same way.
Frequently Asked Questions
Is AI worse at ICD-10 or CPT coding?
In this hand-surgery study, the models were far worse at ICD-10 (23.9% correct) than at CPT (91.5% correct). The result is specific to three procedures and three general-purpose models, so treat it as a signal about where AI coding is weak, not a universal benchmark.
What is the most common AI error in ICD-10 coding, and why does it matter?
In this study the single most common ICD-10 error was incorrect or omitted laterality: the model returned an unspecified-side code, or the wrong side, when the note supported a specific one. Under the FY2026 Guidelines an unspecified-side code should rarely be used, so an AI draft that defaults to it hands the coder a specificity gap to close, not a finished code.
When can I use an unspecified-laterality ICD-10 code?
Under the FY2026 ICD-10-CM Official Guidelines (Section I.B.13), unspecified-side codes should rarely be used, such as when the record does not document the side and it cannot be clarified. If another clinician documented the side you may code from that; if the record conflicts, query the provider.
Does better prompting fix AI laterality errors?
The study found prompt style (zero-shot, one-shot, multishot, chain-of-thought) made no meaningful difference, but a prompt rewritten to emphasize laterality improved accuracy (40%). Even then the authors concluded the models are not ready for independent coding use.
Does this study apply to risk-adjustment or outpatient coding?
No. It tested 90 hand-surgery clinic and operative notes with ChatGPT 3.5, ChatGPT 4.0, and Gemini. It did not test vendor coding tools, ambient scribes, current model versions, or outpatient and risk-adjustment charts.
Sources
- Assessing Large Language Models for Clinical Coding in Hand Surgery: Effect of Note Authorship, Prompt Design, and Diagnosis/Procedure Type — Journal of the American Academy of Orthopaedic Surgeons, Aug 1, 2026
- ICD-10-CM Official Guidelines for Coding and Reporting FY 2026 (Section I.B.13, Laterality) — CMS, Oct 1, 2025
- ICD-10-CM (coding and billing reference page) — CMS, Oct 1, 2025
Related Tools
HCC Buddy Coding Team
Editorial
Every HCC Buddy news article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.
Get CMS Updates in Your Inbox
RADV news, model changes, and coding guidance — within days of CMS publishing, not quarters.
More from CMS Watch
The 2026 ACDIS/AHIMA query standard is final: the outpatient and HCC query rules it sets
September 5, 2026
PolicyAmbient AI scribes omit more than they invent, and an omitted chronic condition is a lost HCC
August 21, 2026
PolicyCPT 2027 sharpens the AI taxonomy. How a tool's role is classified decides how it's coded.
August 14, 2026

