The Vulnerability Management Maturity Model: Where Does Your Program Sit?
A five level evidence model for finding the weakest gate in scope, decisions, delivery, verification, and governance.


A vulnerability management maturity model should expose weak decisions. The problem is not that teams lack a maturity label. It is that a polished scanner can hide improvised scope, ownership, delivery, and verification.
Program maturity is limited by the least reliable gate between a new finding and a verified outcome. If discovery is automated but nobody can name the service owner, the program is not advanced. If prioritization uses threat data but closure means a ticket changed status, the program is not measured. Tools describe capability. Evidence describes maturity.
This model uses five levels across five operating dimensions. It is not a certification. It is a way to find the next constraint, fund one practical upgrade, and prove that the upgrade changed normal work.
Maturity is the quality of the weakest decision gate
Each level requires operating evidence across scope, decisions, delivery, verification, and governance.
What is a vulnerability management maturity model?
A vulnerability management maturity model is a staged test of how consistently an organization finds affected assets, makes risk decisions, delivers treatment, verifies results, and governs exceptions. The output should be a set of evidence backed ratings and a short improvement queue.
SANS introduced its model around five operating areas: prepare, identify, analyze, communicate, and treat. Its core point still holds. Vulnerability management never reaches a finished state. The SANS model published July 6, 2020 gives teams a broad process map. The practical gap is deciding what proof earns a level and what investment should come next.
The volume makes loose scoring expensive. The CVE Program report published August 25, 2026 recorded 15,176 published CVE records in the first quarter and 20,709 in the second. That is 35,885 records in six months. The program notes that quarterly snapshots can change as records are reconciled. A team cannot treat that stream with individual heroics. It needs repeatable gates.
Which dimensions should the model score?
Score five dimensions separately. A single average hides the exact constraint leadership needs to see.
- Scope and evidence: Can the team reconcile approved assets, observed assets, software, exposure, controls, and observation time?
- Decision quality: Does priority use affected state, exploitation evidence, local exposure, business consequence, and uncertainty?
- Delivery: Can work reach an owner with change authority, a safe method, a tested rollback path, and enough capacity?
- Verification: Does fresh evidence prove the vulnerable condition changed, or does the program trust activity records?
- Governance: Are policy clocks, exceptions, approvals, escalation, metrics, and investment choices visible and enforceable?
Use the lowest dimension as the program level. That sounds harsh. It is useful. A level three average built from level four discovery and level one verification still produces unreliable closure. The weakest gate controls the result.
What are the five vulnerability management maturity levels?
Level 0: Improvised
Work begins when a scanner report, audit request, or incident creates pressure. Scope changes by project. Analysts rebuild context manually. Owners are found through chat. Closure depends on a ticket comment. Strong people may still reduce risk, but the result cannot be repeated or forecast.
Evidence of level zero includes unknown coverage, multiple private queues, urgent work with no written reason, and exceptions with no expiry. The first upgrade is not automation. Name the approved scope, one queue, one decision owner, and one definition of verified closure.
Level 1: Repeatable
The team runs a documented vulnerability management process. Asset scope has an owner. Findings enter defined states. Response lanes have clocks. Exceptions record an approver and review date. Teams can follow the same path twice, even if much of the work remains manual.
To prove level one, sample ten findings from the last month. Each should show the affected asset, validation result, decision reason, accountable owner, treatment, and closure evidence. Missing fields count as failures. A diagram is not proof.
Level 2: Measured
The program measures its control loop, not just its backlog. Coverage uses approved scope as the denominator. Freshness is visible. Leaders can see time to validation, time to owner acceptance, blocked days, verified exits, reopened findings, and exception age. Failure reasons are coded well enough to direct investment.
Capacity math becomes unavoidable here. Suppose 200 new findings each need 12 minutes of validation and routing. That is 2,400 minutes, or 40 hours. If the batch arrives twice a month, two analyst weeks are consumed before remediation starts. A measured program either reduces weak inputs, automates stable checks, or funds the work. It does not bury the deficit in overdue counts.
Level 3: Contextual
Decisions use current threat and local environment evidence. CVSS remains a technical baseline, but known exploitation, exploit probability, public reachability, service consequence, and proven controls change the action lane. A risk based process turns those inputs into an owned queue rather than a mysterious score.
FIRST made this direction explicit in its CVSS v4.0 Consumer Implementation Guide, updated January 16, 2026. The guide defines maturity from Base scoring through Threat and Environmental enrichment. It also warns that Base scores alone describe a general scenario, not the deployed environment. This is a recent and important shift: context is now part of the scoring lifecycle, not an optional note after the score.
Level 4: Adaptive
Material evidence changes trigger a controlled response. A new KEV entry, an exposed service, a software change, a failed mitigation test, or an ownership change can reopen and recalculate the case. Stable actions may be automated, but risky changes retain approval, rollout, rollback, and verification gates.
Adaptive does not mean autonomous. It means the system notices meaningful change quickly and routes a reviewable action. Humans still own business consequence, outage choices, risk acceptance, and policy. Remediation, mitigation, and acceptance must remain distinct treatment decisions.
How do you assess maturity without grading yourself kindly?
Use a 30 day operating sample, not interviews alone. Pull ten urgent findings, ten normal findings, five exceptions, and five closed items. Trace each record from scope to fresh verification. Ask for source evidence and timestamps. If the answer exists only because someone reconstructed it for the assessment, score the normal process lower.
Give each dimension a level from zero to four. Then write one sentence for the evidence, one for the failure, and one for the next control. Avoid half points. Precision theater will not improve the decision.
| Test | Evidence to request | Failure signal |
|---|---|---|
| Scope | Approved and observed asset reconciliation | Scanner assets used as the denominator |
| Decision | Inputs, unknowns, reason, and policy lane | Severity sort with private overrides |
| Delivery | Owner acceptance, method, window, rollback | Ticket assigned to asset custody |
| Verification | Fresh observation of expected state | Change completion treated as proof |
| Governance | Policy, exception, escalation, capacity record | Overdue work hidden outside the queue |
Why does maturity matter more in 2026?
Mandiant based its 2026 report, published March 23, 2026, on more than 500,000 hours of incident investigations conducted during 2025. Exploits were the leading initial infection vector for the sixth straight year and accounted for 32 percent of intrusions. Organizations found malicious activity internally in 52 percent of investigations, up from 43 percent a year earlier, yet global median dwell time rose from 11 to 14 days.
Those numbers reward evidence quality, not dashboard volume. Better internal detection means little if an owner cannot act. Faster triage means little if the change cannot be delivered safely. A maturity model should connect every improvement to a reduced decision delay, a smaller unknown set, or stronger proof.
What should a 90 day maturity plan include?
During the first 30 days, establish the baseline sample and select the weakest dimension. Write a measurable target such as “90 percent of urgent findings receive an accepted owner within one business day” or “every closed urgent finding has a fresh verification observation.” Do not choose five targets.
In days 31 through 60, repair the operating path. Define the data contract, owner, failure route, and review cadence. Test on a limited asset group. Deep endpoint context with AI driven analysis can help assemble evidence and draft a decision reason, but the reviewer should see sources, missing information, and the proposed next step before approval.
During days 61 through 90, run the control without special attention. Compare normal results with the baseline. If performance only improves while the project team watches, the process did not mature. Keep the new evidence, update the score, and choose the next lowest dimension.
How should leaders use the maturity score?
Use the score to choose an investment, not to decorate a board report. Show the lowest dimension, the evidence behind it, the decisions it corrupts, the cost of the failure, and the control being tested. Leadership can then choose among reducing scope, adding capacity, changing policy, accepting the exposure, or funding a technical repair.
Keep the report small. One page should show the five dimension scores, changes since the last sample, the active constraint, its owner, the target evidence, and the date of the next normal cycle. A long list of planned capabilities makes weak accountability look like a roadmap.
Tie funding to an observable operating result. If the request is for better asset identity, the target might be to reconcile 98 percent of approved assets to fresh observations and assign a failure reason to the rest. If the request is for automation, the target might be to cut median validation effort from 12 minutes to four while keeping reviewer correction below an agreed threshold. A tool launch is an activity. A changed control is an outcome.
Which maturity assessment mistakes produce false confidence?
The first mistake is scoring policy instead of behavior. A written rule can exist while normal work happens in spreadsheets and chat. Sample operating records. The second is averaging away a failed gate. Report each dimension and let the lowest one set the level.
Another mistake is awarding points for product features. A platform may support asset reconciliation, threat enrichment, approval, and verification. The program earns credit only when those capabilities are configured, owned, used, and tested. Shelfware has no maturity value.
Teams also confuse fewer findings with better control. A smaller queue may reflect good validation, or it may reflect missing scope and suppressed exceptions. Demand the denominator and the exit reason. Finally, do not make level four the automatic goal. Adaptive triggers on noisy evidence can multiply bad decisions. Raise a dimension when its current failure costs more than the next level will cost to operate.
Frequently asked questions about vulnerability management maturity
What is a good vulnerability management maturity level?
The right level is the lowest one that supports the organization's risk, scale, and obligations. Most teams gain more from making level two reliable across all dimensions than claiming level four automation over weak scope.
How often should a maturity assessment run?
Run a full evidence sample every six months and after a material change in scope, tooling, policy, or ownership. Review the active improvement metric monthly.
Can a tool move the program up a maturity level?
Only when it changes normal operating evidence. A tool may improve collection, context, routing, or verification. The level rises after the team proves the related decision works consistently.
Should every dimension reach level four?
No. Automation has cost and failure modes. Raise a dimension only when the next level improves decision quality, speed, or control enough to justify the operating burden.
The executive takeaway
Stop asking whether the vulnerability program is mature. Ask which gate fails under ordinary pressure. Sample 30 records, score the five dimensions, and fund the lowest one. Require operating evidence after 90 days. A lower honest score with one repaired constraint is more valuable than an advanced label built on a clean dashboard.
Put more evidence behind vulnerability decisions
Artemes AI combines endpoint telemetry, sourced vulnerability intelligence, and analysis with practitioner review so teams can examine the evidence, missing context, and recommended next step together. We are accepting early access requests now.

Alex Gibson
Alex writes about configuration drift, operational security evidence, endpoint telemetry, triage supported by AI, and the practical work of turning signals into better remediation decisions.
Related Reading
Get articles like this in your inbox.
Security research and occasional Artemes AI product updates.

