Updates
Mark Cuban IntegrationDPC Summit 2026Webinar with DPC AllianceMIPS on ElationSpruce PartnershipHint Partnership
HealthCompiler
Platform
Solutions
Who We Serve
Resources
APEX
Sign upRequest a Demo
K
HealthCompiler

+1 408 883 7902

Health Compiler Inc.

2261 Market Street #4632

San Francisco, CA 94114

All Rights Reserved | Health Compiler Inc. © 2026

Made with ❤️ in San Francisco

Privacy Policy|Terms of Service|Security

QUICK LINKS

  • Home
  • Integrations
  • Forward Deployed Engineering
  • Employer Analytics
  • Health Outcomes
  • AI Call Triaging

RESOURCES

  • APEX
  • Blogs
  • About Us
  • FAQs
  • Whitepaper

REACH OUT

  • Tech Support
  • Success Stories
  • Sales
  • Careers
  • Contact Us
HIPAA Compliant and AICPA SOC 2 Type II certifiedFree Market Medical AssociationDPC AllianceChrome StoreGoogle Cloud Marketplace
September 23, 2026
5 min read
RAF Scores Explained: What Accuracy Actually Requires in 2026

RAF Scores Explained: What Accuracy Actually Requires in 2026

A RAF score estimates what it should cost to care for a Medicare Advantage member, based on demographic factors and the health conditions on file. It is easy to treat that number as a revenue lever. It works better as a description of how sick your members actually are.

The difference matters, because a score is only worth what the evidence behind it is worth. A higher number built on thin notes is not an asset. It is a liability sitting on your books until somebody asks to see the chart.

Two developments in 2026 make that point sharper than it has been in years. The 2024 CMS-HCC model is now fully in force, and CMS has rebuilt its audit operation at a scale the industry has not seen before.

What the score measures

RAF stands for Risk Adjustment Factor. CMS uses risk adjustment because some member groups cost far more to care for than others, and plans should not be punished for enrolling sicker people. A member's score comes from demographic factors plus qualifying conditions, submitted through accepted encounter and diagnosis data.

A score near 1.0 represents average expected cost under the model. Above 1.0 means higher expected cost, below means lower. The payment effect is never a fixed dollar figure, because payment also turns on county benchmarks, plan bids, quality bonuses, and payment-year rules.

Underneath, CMS-HCC models sort selected ICD-10-CM diagnoses into Hierarchical Condition Categories. The hierarchy reflects severity, and it prevents a plan being paid twice for two closely linked signs of the same underlying disease. Demographic factors and condition weights then combine into the final score.

What the 2024 model changed

The 2024 CMS-HCC model, known across the industry as V28, replaced the 2020 model over a three-year phase-in.

The restructuring was substantial. CMS rebuilt the condition categories around ICD-10 rather than the older ICD-9 system, and recalibrated on newer data, moving from 2014 diagnoses and 2015 spending to 2018 diagnoses and 2019 spending. The stated aim, set out in the CY2024 Rate Announcement, was to better reflect current disease patterns, treatment methods and costs, and coding practice.

The model now holds 115 payment HCCs against 86 in the previous 2020 model, plus a further 151 categories outside payment altogether. Depression and diabetes saw the largest revisions. CMS moved certain depression codes it judged weaker at predicting cost into non-payment categories, keeping the remaining 350. More than 300 diabetes codes remain.

That phase-in is now finished. For 2026, CMS is calculating 100% of risk scores using the 2024 model alone, so any comparison you run against prior years is comparing two different rulebooks.

If you relied on static spreadsheets or vendor black boxes through the switch, explaining why a given score moved is now harder than it should be.

What CMS accepts as proof

This is where many programs are weaker than they think, and CMS states the standard plainly in its RADV audit instructions.

A valid medical record is a legibly signed and dated record showing diagnoses and services from a face-to-face visit, conducted by an appropriately credentialed provider, within the data collection period. That period is the calendar year before the payment year.

Several specifics follow. Records from hospital inpatient units, outpatient sites, and physician office visits are accepted. Hospice records, home health records, lab-only records, superbills, and any encounter that was not face-to-face are not. The signature must belong to the credentialed provider who actually saw the patient. Where a signature is missing, an attestation may cover physician and outpatient records, but never inpatient hospital records.

One consequence catches programs out every year. The collection period resets annually, so a chronic condition written up last year does not support this year's score on its own. It has to be assessed and documented again. A problem list is a prompt to take another look, not evidence.

Why the audit math changed

In May 2025, CMS announced a substantial expansion of Risk Adjustment Data Validation auditing. The specifics change the exposure calculation for every plan.

CMS moved from auditing roughly 60 plans a year to auditing all eligible MA contracts for each payment year in newly initiated audits. That works out to around 550 plans annually. Records pulled per plan rose from 35 to between 35 and 200, depending on plan size.

The agency grew its medical coding team from 40 people to roughly 2,000 and said it would deploy advanced systems to flag unsupported diagnoses. It set a target of completing audits for payment years 2018 through 2024 by early 2026, and findings can be extrapolated where the RADV final rule allows.

For most contracts, being audited is no longer a question of probability. Thin evidence, copied-forward problem lists, coding more specific than the notes support, and loosely governed suspecting are all considerably more likely to surface.

Where accuracy actually breaks down

Three problems cause most of the damage, and none of them are coding problems.

The data is scattered. A member's clinical story sits across primary care notes, specialist visits, claims, pharmacy activity, lab results, hospital feeds, and health information exchanges, and no single source tells the whole story. Without solid identity matching and a linked history, teams either miss evidence or attach it to the wrong member.

Care and coding run on different clocks. Care happens during a visit, while claims and encounter data arrive weeks or months later. By the time a year-end review finds a gap, the provider may have no clinical reason to see that member again. This is why support delivered before or during the visit outperforms chart review after the fact.

And the work gets treated as a project rather than a process. A year-end push can only recover what the calendar still allows. Everything else has already expired.

What a working program looks like

The groups that get this right treat risk adjustment as a running clinical data process. In practice that means five things connected end to end.

Start with one linked member record, pulling together eligibility, claims, encounters, EHR data, pharmacy, and lab results. Resolve identities, standardize terminology, and retain source lineage, so you know where every element came from and when it arrived.

Calculate with model logic you can see into. Map diagnoses to the correct payment-year model and apply hierarchy and interaction rules. Every component should be explainable, so a team can move from a population trend down to the member, the condition, the visit, and the source document.

Rank the cases worth a clinician's time by confidence, documentation status, timing, and likely care impact. Suppress stale or low-value alerts so providers are not buried in noise.

Put the work inside clinical workflow rather than alongside it, using short evidence-linked prompts before or during the visit. The provider keeps clinical judgment and decides whether the condition is present, addressed, and properly written up.

Finally, keep an audit trail recording source, model version, evidence, user action, outcome, and submission status for every condition. That is what turns a defensible claim into a provable one.

Average RAF alone will not tell you whether any of this is working. Track risk-score trend by cohort and provider, and recapture rate for known conditions broken out by HCC. Watch provider acceptance and rejection rates, suspect precision, and the time from finding a case to closing it. Those measures separate real improvement from score inflation.

Where AI helps and where it stops

AI is useful here. It can organize large volumes of structured and unstructured data, surface supporting evidence, summarize a member's history, and rank review queues. That lets teams concentrate on the members most likely to benefit.

What it should not do is decide a diagnosis, invent specificity, or turn a weak signal into a submitted code. CMS validates a diagnosis against a signed record of a face-to-face visit by a credentialed provider. The clinical judgment has to be human and written up as such. Every proposed condition should be explainable, linked to evidence, reviewed under your own policy, and confirmed by a qualified clinician or coder where required.

How Health Compiler fits

Health Compiler connects healthcare data and turns it into governed, usable intelligence. For Medicare Advantage plans, ACOs, provider groups, and value-based care organizations, that means linked member views across claims, EHR, lab, pharmacy and eligibility data. It also means CMS-HCC model processing, recapture of known conditions alongside clinically backed suspect finding, provider worklists shaped around existing workflow, and evidence lineage that stands up to review.

Health Compiler's integrations connect EHR, claims, pharmacy, lab, eligibility, and operational data. Its Health Outcomes platform helps care teams act on linked risk signals, and Forward Deployed Engineering adapts the data and workflow layer to the systems each organization actually runs.

Disclaimer: This article is for general information only and is not legal, compliance, coding, or financial advice. Risk adjustment rules change; verify all requirements against current CMS guidance and consult your own advisors before acting. Health Compiler is not affiliated with or endorsed by CMS. Diagnosis and documentation decisions remain the responsibility of the treating clinician and the submitting organization.

FAQs

What does a RAF score of 1.0 mean? It represents average expected cost under the model in use. It is not a fixed payment amount, and the real financial effect varies by contract and payment method.

Does every diagnosis raise RAF? No. Only diagnoses that map to the model in use and meet submission and documentation rules can move the score. Hierarchies can also let a more severe condition override a related, milder one.

Do chronic conditions carry forward automatically? No. A condition must be assessed and documented within the data collection period, which is the calendar year before the payment year. Past data is best used to prompt a fresh clinical look.

What documentation does CMS accept? A legibly signed and dated medical record showing diagnoses and services from a face-to-face visit by an appropriately credentialed provider, within the data collection period. Hospice, home health, lab-only records, and superbills are not accepted.

Can AI assign HCC diagnoses? AI can find evidence and suggest conditions for review. The diagnosis and the record behind it must rest on the patient's chart and proper professional judgment.

What changed with V28? The 2024 CMS-HCC model rebuilt its condition categories around ICD-10 and recalibrated on 2018 diagnoses and 2019 spending. It holds 115 payment HCCs against 86 in the 2020 model, with major changes to depression and diabetes. For 2026, CMS calculates 100% of risk scores using this model.

LinkedInTwitter / XFacebookPinterestGoogle+Email
Read More Articles