What AI Can and Can't Do in HCC Risk Adjustment Coding Right Now
A risk adjustment QA perspective on AI in HCC coding, what works, what fails, and what every coder needs to know to stay ahead in 2026.
By the HCC Buddy Coding Team
Updated: September 5, 2026

The Conversation Nobody in Risk Adjustment Is Having Honestly
Every week, coding teams see another headline: "AI Processes Thousands of Charts Per Day." "AI-Driven Coding Accuracy Hits 95%." "The Future of Medical Coding Is Autonomous."
And every week, QA teams still pull charts where the AI-suggested codes would have failed a RADV audit.
HCC coding QA exists to catch errors before they become compliance problems. AI coding tools have moved from "interesting experiment" to "thing coders are being told to use" in less than 18 months. Some of what these tools do is genuinely impressive. Some of it is dangerous. And the conversation happening in our industry right now is not distinguishing between the two.
This is not an anti-AI article. The point is to be specific about what works, what does not, and what recent Medicare Advantage enforcement actions should teach risk adjustment coders about source and documentation review.
What AI Actually Gets Right in HCC Coding
Credit where it is earned: there are things AI coding tools do well, and pretending otherwise would be dishonest.
Speed on straightforward lookups. If a provider documents "Type 2 diabetes mellitus with diabetic chronic kidney disease, stage 3," an AI tool can map that to E11.22 and show you it hits HCC 37 (Diabetes with Chronic Complications) under V28 almost instantly. For clean, well-documented, single-condition encounters, AI is faster than any human.
Pattern recognition across large chart volumes. AI is good at flagging charts where a medication list suggests a diagnosis that was not captured. Patient is on insulin, metformin, and lisinopril but the encounter only documents hypertension? The tool flags the missing diabetes. This is real value, it catches the low-hanging fruit that human coders sometimes miss at hour six of a ten-hour shift.
V28 mapping updates. The source-checked comparison contains 8,019 unique payable-code mappings in CY2026 V28 and 9,927 in the historical PY2025 V24 reference. Of the union, 2,267 appear only in historical V24 and 359 only in V28. The safe workflow is to refresh from the official source and verify the exact code, model, and payment year rather than rely on an old printed crosswalk.
Volume processing for initial screening. For large retrospective chart reviews, AI can do the first pass, identifying charts that likely contain capturable HCCs, and send only the flagged charts to human coders. This triage function is legitimate. It lets your team focus their clinical judgment on charts that need it rather than spending time on charts with nothing to capture.
These are not trivial benefits. For coding operations processing thousands of charts, the productivity gain on routine work is real.
Where AI Falls Apart, and Where It Gets Dangerous
Here is where the real risk starts.
Documentation support still needs human review. MEAT, Monitoring, Evaluation, Assessment, and Treatment, is a common documentation-review mnemonic. It does not replace the official ICD-10-CM, CMS, encounter, program, or payer rules that determine whether a diagnosis is reportable.
AI tools can find a diagnosis mentioned in a chart. What they cannot reliably do is determine whether the provider actually monitored, evaluated, assessed, or treated that condition during that specific visit. Did the provider review the latest A1C and adjust the diabetes management plan? Or did the problem list just carry forward "Type 2 diabetes" from the last visit with no clinical engagement?
This distinction matters in risk adjustment. In January 2026, Kaiser affiliates agreed to pay $556 million to resolve False Claims Act allegations involving Medicare Advantage diagnosis submissions. According to the Department of Justice settlement announcement, the United States alleged that, from 2009 to 2018, Kaiser used data mining and post-visit queries to urge physicians to add diagnoses through medical-record addenda. The settlement resolved allegations only, with no determination of liability.
That case is a caution against treating software suggestions as coding conclusions. A tool can surface a possibility. The applicable source, medical record, and human review still have to support the submitted diagnosis.
Multi-condition hierarchy logic breaks AI. V28 has 115 HCC categories arranged in hierarchical groups. When a patient has multiple related conditions, the hierarchy determines which code trumps which. A patient with both diabetes with chronic complications (HCC 37) and diabetes without complication (HCC 38): the more severe HCC 37 trumps HCC 38, so only HCC 37 is counted. But the clinical documentation has to support the higher-severity code specifically.
AI tools regularly suggest the higher-paying code in a hierarchy without evaluating whether the documentation actually supports that level of severity. QA teams have seen AI suggest Stage 4 pressure ulcer codes when the clinical note describes a Stage 2. The tool read "pressure ulcer" and picked the code with the highest RAF weight. That is not coding. That is upcoding.
Specificity gaps create audit exposure. V28 demands more clinical specificity than V24 ever did. "Heart failure" is not enough, you need systolic versus diastolic, the ejection fraction, the NYHA class. "CKD" is not enough, you need the exact stage. AI tools often default to the most general code when documentation is ambiguous, or worse, infer specificity that is not explicitly documented.
In a RADV audit, the question is simple: does the medical record support this exact code? If the AI suggested E11.65 (Type 2 diabetes with hyperglycemia) but the note only says "diabetes, sugars running high" without a documented glucose reading or A1C showing hyperglycemia, that code gets rejected. Your RAF score gets clawed back. At scale, those clawbacks add up to millions.
Context and clinical judgment cannot be automated. Here is a common example. A chart documented "history of breast cancer, currently on tamoxifen, no evidence of recurrent disease." An AI tool flagged this for HCC capture as an active malignancy. But "history of" with "no evidence of recurrent disease" is not an active cancer, it is a personal history code (Z85.3), which does not map to an HCC. A coder with clinical judgment reads that documentation and immediately knows the difference. The AI read "breast cancer" and "tamoxifen" and drew the wrong conclusion.
This is not a rare edge case. Cancer history versus active cancer. Resolved versus chronic conditions. Acute exacerbation versus stable chronic disease. These distinctions drive HCC capture decisions dozens of times per day, and AI gets them wrong often enough that you cannot trust it without human review.
The Enforcement Landscape Has Changed, and AI Is Not Ready for It
Here is what is happening on the compliance side, because this is the context that makes the AI accuracy problem urgent rather than academic.
CMS continues to conduct contract-specific RADV audits. The current CMS MA RADV program page explains that CMS checks whether diagnoses submitted for risk adjustment are supported in enrollees' medical records and may recover overpayments when they are not.
The Kaiser settlement is one recent enforcement example. The $556 million settlement resolved allegations only and did not determine liability. DOJ says two former Kaiser employees, Ronda Osinek and James M. Taylor, M.D., brought the qui tam cases; the combined relator share was $95 million.
In this environment, "the AI suggested it" is not a defense. CMS does not care whether a human coder or an AI tool selected the diagnosis code. The question is whether the medical record supports it. Period.
The industry is shifting from what Fierce Healthcare called "coding intensity", capture as many codes as possible, to "defensible accuracy", can you prove every single code you submitted? AI tools that were designed for the coding intensity era are liabilities in the defensible accuracy era.
What Smart Coders Are Actually Doing With AI
Coders using AI effectively are not letting it code for them. They are using it as a pre-screening layer and then applying their own clinical judgment to every single code before it goes out.
Here is what that workflow looks like in practice:
AI does the first pass. The tool scans the chart and flags potential HCC-capturable diagnoses. This saves the coder from reading every line of a 40-page chart looking for conditions that may or may not be there.
The coder validates every flag. For each suggested code, review the full record, provider diagnostic statement, official classification, and applicable encounter and program rules. MEAT may organize questions, but it does not validate the code.
The coder catches what AI misses. AI tools miss diagnoses that are documented narratively rather than in a structured problem list. A provider who writes "patient's cognitive decline has worsened, now requiring 24-hour supervision, family considering memory care placement" has documented dementia progression, but the AI might not flag it if the word "dementia" does not appear.
The coder rejects what AI gets wrong. Every AI suggestion that does not pass the coder's clinical judgment gets rejected. No exceptions. The coder is the last line of defense before that code hits CMS.
This is the workflow that survives a RADV audit. The AI makes the coder faster. The coder makes the AI accurate. Remove either one and you have a problem.
The Real Question Is Not "Will AI Replace Me?", It Is "Am I Using the Right Tools?"
The anxiety is understandable. When you read that AI processes thousands of charts per day and you process 50, the math feels threatening. But here is what those headlines leave out.
Automated suggestions still need qualified human review against the medical record and the rules that apply to the submission. Throughput before that review is not a finished coding result.
The coders who are genuinely at risk are the ones using outdated tools and outdated workflows, not because AI will take their jobs, but because they cannot keep up with the accuracy demands of V28, the speed of quarterly RADV audits, and the documentation specificity that defensible accuracy requires.
If you are still coding from printed reference guides, manually cross-referencing V24 to V28 mappings in a spreadsheet, or switching between six browser tabs to look up codes, check HCC mapping, calculate RAF scores, and verify hierarchies, you are spending your time on things a tool should handle so you can spend your judgment on things only a human can.
The coders who will thrive in 2026 and beyond are the ones who pair their clinical judgment with tools purpose-built for risk adjustment, tools that give them instant V28 mapping, RAF calculations, drug-to-diagnosis cross-references, and hierarchy checks so they can focus their expertise on the documentation review, specificity decisions, and interpretation of the record that still require coder judgment.
That is why HCC Buddy exists. It does not replace coders. It makes source-labeled lookup and comparison faster so coders have more bandwidth for work that requires judgment. The Chrome extension keeps HCC lookups and drug references beside browser-based chart work. Member scoring stays in the RAF tool, and HCC Buddy checks updates against official files before showing them.
What You Should Do Right Now
Whether you use HCC Buddy or not, here is what the current landscape demands from every risk adjustment coder.
Audit your own work against the actual rules. Pull a sample of coded charts and verify every submitted diagnosis against the full record, official classification, eligible encounter and provider source, and applicable program instructions. Use MEAT only as an optional review aid.
Learn the current source boundary. CY2026 non-PACE risk scores use V28 at full weight, while PACE uses a separate blend. The checked comparison has 2,267 codes found only in the historical PY2025 V24 mapping and 359 found only in current V28. Verify the exact code instead of carrying a historical result forward.
Understand what your AI tools are actually doing. If your organization is pushing automated coding, ask specific questions. What is the false positive rate? How does it check documentation support against the applicable official rules? Can it distinguish active conditions from historical ones? If nobody can answer those questions, the tool is a compliance risk.
Do not let anyone, human or AI, pressure you to submit codes you cannot defend. DOJ says two former employees brought the Kaiser qui tam cases and that their combined relator share was $95 million. The practical lesson is simpler: do not let a productivity metric or software suggestion override the applicable source and the medical record.
The era of coding intensity is over. The era of defensible accuracy is here. The coders who understand that distinction, and have the tools to operate within it, are the ones who will build long, successful careers in risk adjustment.
Related Tools
ICD-10 Encoder
See CY2026 V28, historical PY2025 V24, and source status before using a mapping or factor.
RAF Calculator
CMS-HCC V28 Payment Year 2026 scoring with complete member context. No score is shown unless the required source and calculation checks pass.
Chrome Extension
Use source-labeled HCC lookups and drug references beside browser-based chart work.
HCC Buddy Coding Team
Editorial
Every HCC Buddy article is checked against the current CMS-HCC model and the active FY ICD-10-CM tabular release before it publishes.
Get HCC Coding Tips in Your Inbox
Join our newsletter for coding tips, guideline updates, and tool announcements.




