Most health plans I talk to are running their next Risk Adjustment (RA) vendor evaluation on the same framework they used two years ago. Risk Adjustment Factor (RAF) lift per dollar, contract coverage across the book, coder productivity, chart retrieval turnaround, Hierarchical Condition Category (HCC) counts against a benchmark. That framework measured what mattered under the retrospective payment model most of the industry ran on for a decade. It stopped measuring what determines RA outcomes in the environment CMS built over the last 18 months, and plans that keep evaluating against it will keep picking vendors that underperform under the new conditions. That is not a comfortable position to defend to a CFO who signed off on last cycle’s vendor selection using the same criteria the plan is about to run again, but it is where the argument lands when you look at what V28, chart review exclusion, and Risk Adjustment Data Validation (RADV) expansion did to the assumptions underneath the retrospective framework.
This is Article 1 of a three-part series called Risk Adjustment in the New Regulatory Environment. It covers the plan-side buying criteria that stopped fitting. Article 2 moves to the risk adjustment vendor side of the market and what the shift means for those businesses. Article 3 covers the organizational separation between Risk Adjustment and Quality Improvement (QI), and why the environment CMS built makes that separation structurally unsustainable.
Why the Retrospective-Era Evaluation Framework Worked for a Decade
The retrospective RA evaluation framework was built for an environment where the primary revenue driver was catching diagnoses the visit missed. Chart review, sweeps, and analytics-driven suspecting satisfied much of this, and every criterion on the vendor evaluation measured how effectively a vendor could run one or more of them. RAF lift per dollar rewarded vendors who moved the most incremental RAF for the least fee. Contract coverage rewarded vendors who could operate across the entire book. Coder productivity and chart retrieval turnaround rewarded vendors who could staff and scale retrospective operations. And HCC counts against benchmark rewarded vendors who were thorough. The framework measured retrospective throughput because retrospective throughput produced most of the RAF.
Under Version 24 of the CMS Hierarchical Condition Category (CMS-HCC) model, retrospective recovery closed most of the gap between what the visit captured and what the plan submitted. Audit frequency was low enough that the exposure created by aggressive retrospective work rarely surfaced at the enterprise level, and the Medicare Payment Advisory Commission (MedPAC) projected MA overpayments at roughly $84 billion for 2025 largely from that gap. Both plans and vendors optimized against what CMS was paying, and the framework rewarded the mechanics of running that operation well. And for a decade, it did.
How the Current Regulatory Cycle Broke That Framework
The three CMS actions of the last 18 months changed what vendors have to be good at. Version 28 of the CMS-HCC model drives 100% of risk score calculations as of Calendar Year (CY) 2026, retiring roughly 2,000 payment-eligible diagnosis codes. Chart review exclusion, finalized in the CY 2027 Final Rate Announcement, closes off risk-adjusted payment on diagnoses from unlinked chart reviews, a channel nearly 58% of MA contracts used in 2022. And RADV expansion moved audit from roughly 60 contracts a year to every eligible contract annually, with sample sizes up to 200 records per contract. Enforcement arrived alongside: Kaiser Permanente at $556 million in January 2026, Aetna at $117.7 million in March, Elevance at $342 million paid on a $935 million reserve in May, and provider group The Villages Health System at $541.5 million in August tied to invalid diagnoses submitted through delegated risk arrangements. The first MA-specific OIG compliance guidance since 1999 landed in February 2026 flagging add-only chart reviews and in-home health risk assessments.
Together, those actions reprice what a vendor has to do to produce durable RA outcomes for a health plan. The retrospective evaluation measured throughput on channels CMS is either shrinking or compressing. What it doesn’t measure is what actually determines whether the RAF a vendor helps a plan capture holds up under the audit cadence CMS locked in this year, and whether the plan’s overall RA operating model is being pushed in the direction the current regulatory cycle rewards.
The Four Criteria That Fit This Environment
The evaluation criteria that fit the current environment measure four distinct properties. They apply to different vendor categories in different ways, and taken together they form the framework that produces defensible RAF into CY 2028.
Audit defensibility. This criterion applies most to medical coding vendors, and to a lesser extent to submission platform vendors. Audit defensibility is not the same thing as RADV support services, and confusing the two is one of the more common evaluation errors I see. Audit defensibility is a property of how the coding work gets done in the first place, upstream of any audit event. What plans are actually evaluating on this criterion is the depth of the coding team’s regulatory literacy, coding guideline expertise, and experience defending prior coding decisions under audit scrutiny. Coders who understand CMS guidance well enough to push back on the plan’s own coding assumptions, educate plan leadership on where documentation integrity is thin, and increase the rigor of the plan’s coding decisions produce work that holds up. Technology matters here too, whether that shows up as workflow, platform, or AI-enabled coding review, but technology sits alongside the coding team’s judgment rather than replacing it. Plans should be evaluating the vendor’s coder training, the vendor’s own audit history on prior contracts, and the vendor’s willingness to say no to codes the plan wants captured.
Encounter integration depth. This criterion applies to a narrower set of vendors: coding companies that have built EMR integration, provider-facing coding tools, and clinician workflow integration into their delivery model. It does not apply to every coding vendor. The RA workflow CMS is not compressing is the one that captures diagnoses at the visit itself, and vendors that can operate inside the encounter through EMR access, clinician-facing suspecting, and provider workflow integration are the ones working the channel the current conditions leave open. What plans should be evaluating here is the depth and quality of the vendor’s provider-side deployments, not the vendor’s marketing claims about encounter capabilities. Two questions matter more than the rest: which EMRs the vendor is actually integrated with beyond generic Health Level Seven International (HL7) connectivity, and how the vendor’s tooling behaves inside the clinician’s workflow rather than around it. Vendors that can produce named provider deployments with retained utilization data have the depth.
Provider and delegated risk enablement. This criterion applies to a different vendor category entirely: value-based care enablement companies, wraparound network operators, and third-party groups that take upside and downside risk on defined member populations or geographies. These are not coding vendors, and evaluating them on a coding scorecard misses what they actually do. The RAF capture happening on the encounter-level side of the market is largely happening through arrangements these vendors enable, and plans that want to migrate coding upstream into the provider relationship are doing that through delegated risk deals these vendors structure and support. What plans should be evaluating here is the depth of the vendor’s risk-taking arrangement, the specificity of the coding accountability written into the delegated deal, the analytics infrastructure the vendor provides to the participating providers, and the vendor’s actual willingness to take defined downside exposure alongside the upside. Vendors that structure delegated risk deals with coding accountability built into the contract and defined downside exposure are enabling the operating model the environment rewards.
RADV specialist capability. This one is different from the first three, because it is not a criterion plans typically apply to their coding vendor at all. RADV specialist work is a distinct capability performed by a small set of specialty firms. Not every retrospective coding vendor has the audit response expertise, regulatory depth, or extrapolation methodology to run a RADV response engagement at any scale, and many plans deliberately separate the two decisions to avoid the fox-guarding-the-henhouse dynamic of hiring the same vendor to code and defend the code. Under the old framework, RADV was infrequent enough that plans could treat this as an event-based decision. Under an all-contract annual audit cadence, RADV specialist selection has to become a standing decision alongside the coding vendor evaluation, and the two evaluations should stay structurally separate. Plans that fold RADV response into their coding vendor scorecard as a service line get shallow RADV capability at best and a conflict of interest at worst.
How to Assess Your Current Vendor Stack Against the New Criteria
The four criteria above are useful the moment a plan runs its next full RA vendor evaluation, but most plans reading this aren’t running a clean-sheet evaluation next quarter. Instead, they're managing an existing vendor stack under contracts written to the old framework, with renewal decisions coming due on staggered timelines. The practical starting point isn’t the next evaluation. It’s taking stock of what the plan already has against the criteria that matters most.
A quick way to take stock is to score the current vendor stack against the four criteria the way you would score a new vendor evaluation. That exercise usually surfaces two or three specific gaps that matter more than the rest. A plan that re-evaluates its current coding vendor against audit defensibility rather than RAF lift per dollar can raise what it finds as a performance conversation inside the existing contract, before the renewal window forces a yes-or-no decision. A plan that scores its current coding vendor against encounter integration depth can see how far that vendor is actually operating inside the clinical workflow versus around it, and whether the gap is one the vendor can close through delivery changes under the current contract or one that calls for a different vendor type at renewal. The exercise is less about producing a replacement list and more about surfacing the specific gaps that will cost the plan the most if they go unaddressed through the next contract cycle.
CY 2027 vendor decisions are being made right now. Plans that wait until CY 2028 to update the evaluation framework will spend CY 2027 operating against criteria that no longer fit, and the vendor commitments they lock in during this cycle will run against a regulatory backdrop that has already moved. Updating the criteria now is what positions a plan’s vendor stack to hold up under the audit cadence CMS locked in this year.
What Comes Next in This Series
Article 2 takes this to the risk adjustment vendor side of the market and what the shift means for how those businesses are built, run, and valued. Article 3 covers the organizational question underneath both: the separation between Risk Adjustment and Quality Improvement inside plans, and why the environment CMS built makes that separation structurally unsustainable.
Ryan Peterson is the Founder and Principal of Upward Growth, a health plan market advisory firm that works with health tech vendors, investors, provider organizations, and management consultancies to build strategy around how health plans actually buy, operate, and make decisions. He has 15 years of health tech experience across go-to-market and strategy roles serving the health plan market. He also publishes the weekly Upward Growth newsletter at healthplanweekly.com, covering payor market trends, health tech go-to-market strategy, and CMS regulatory developments for a subscriber base of 7,000+ healthcare executives and investors.