top of page

TQVM - Total Quality Vulnerability Management

  • 11 hours ago
  • 9 min read

VM is a process problem.  AI ended the grace period.


WHAT THIS IS ABOUT

Rework is the most expensive way to get anything done.  Manufacturing learned that decades ago and stopped depending on inspection to produce quality.  Vulnerability management never made the move - it still scans, counts defects, and routes them to whoever will take them.


VM must be run like a chip factory; chasing findings will never suffice again.

This paper covers vulnerability management because that is where the gap is widest, but the approach is not specific to it.  For example, NIST asks software producers to "address the root causes of vulnerabilities to prevent recurrences" rather than find and fix instances after release, and DevSecOps is that principle in practice.7  What follows applies the same discipline to the infrastructure estate that is already running.

 

Cease dependence on inspection to achieve quality.  Eliminate the need for inspection on a mass basis by building quality into the product in the first place.

W. Edwards Deming, Out of the Crisis, 19821


Open your last set of Vulnerability Management (VM) metrics.  If they are dominated by vulnerability counts, severities, defect aging, and rack-and-stacks by business unit or application group, you are measuring activity, not risk.  Defect counts do not tell you whether the estate is under control.  They tell you how verbose your scanner was last month.


The short version.  Vulnerability state is determined by what is installed, how it is configured, how often it is updated, and whether it is still supported.  Every one of those is a process, and every one of them is run by infrastructure and application teams.  Cybersecurity is accountable for the outcome and executes none of them - which is why the answer is not better findings.  It is better strategy.


Cybersecurity must keep the accountability.  What has to change is what it demands.  Stop counting what was closed and start measuring the ITAM and ITSM processes that truly control vulnerability risk.


CYBER MUST DEMAND QUALITY - NOT REWORK.

 

WHY THIS CAN'T WAIT

Rework was always the expensive way.  Now it is also too slow to matter.  (Paper 5 - What Changed.)  Mandiant's mean time to exploit fell from sixty-three days in 2018-19 to an estimated minus seven in 2025 - a figure driven by edge and core network devices, which is precisely the part of the estate least likely to sit inside a standing service.2  On an independent dataset, VulnCheck found that of the 884 vulnerabilities first observed under exploitation in 2025, twenty-nine percent were already being exploited on or before the day the CVE was published.3  While the numbers can be debated, the trend is unequivocal - vulnerabilities cannot be chased.


Be clear about what this proves.  No patch cycle beats an exploit that arrives first - that gap is what mitigation is for.  What the closing window does prove is that negotiating remediation one instance at a time is finished.


Cloud and ephemeral infrastructure only prove the need to evolve.  Rebuilding rather than repairing means managing one image instead of countless servers - but instances still have to be governed, so the problem moves from chasing hundreds of assets to establishing one solid process.


The alternative already runs at scale.  Large infrastructure estates execute regular remediation across hundreds of thousands of assets - tens of millions of changes a year - without a business-impacting incident attributable to the service.4  Mature environments applying regular, incremental changes to well-supported assets rarely see a hiccup.  Immature and ad hoc processes present risk with every change they make.


The gap between those two is not technology or magic.  It is process and rigor.  ITAM defines the estate.  ITSM executes the remediation.  Vulnerability management is the seam between them, and none of it is improved by ranking findings better, by pressuring stakeholders who cannot act, or by buying a tool that does both more expensively.

 

WHAT INSPECTION IS ACTUALLY FOR

Manufacturing still inspects.1  Inspection just stopped being how quality was produced - it became how the process is verified.  Apply that lens and the familiar VM objects change meaning:

  • A finding inside its committed time is work in process.  It is work the service has not reached, not work it failed to do.

  • A finding past its committed time is a defect - and it is a defect of the service, not of the asset.

  • A remediation SLA is not a deadline you impose on somebody.  It is the cycle time your service is capable of - and if you have to impose it, you do not have a service.

  • Service maturity is the share of the estate whose changes are executed by a repeatable process rather than by negotiation.


Vulnerability state is derived, not observed.  If an asset is not properly managed - no lifecycle discipline, no configuration hardening, no reliable update cadence - you already know it carries critical vulnerabilities.  You do not need a scanner to tell you that, and if the scanner is the first place you find out, that is the finding.

The scanner should be verifying a prediction, not producing a discovery.  Its job is to show where your ITSM processes are failing and your ITAM is wrong.  That is process control, not risk management, and it never has been.

 

THE CHANGE: MANAGE THE ESTATE, NOT THE FINDINGS

Manage and protect your estate.  Stop chasing vulnerabilities.


Nothing here is new - it's simply a change in focus and a change in thinking.

  1. Own everything.  (Paper 1 - Zero Vulnerability Framework.)  By definition, every vulnerability lives on an asset, and every asset must have an owner, and services that support it.  If those owners are unknown today, that is not an objection to the model. 


    That is your starting point.  Call in the ITAM squad and get to work.  And apply the obvious corollary: the most efficient way to eliminate vulnerabilities is to eliminate unused and underused assets.  It is the only action that permanently removes current and future vulnerabilities alike.


  2. Organize work into actions performed by services.  (Paper 2 - Service-Oriented Vulnerability Management.)  The unit of work is the action, not the vulnerability - one monthly patch application closes a hundred findings at once, and one core OS service applies it across a thousand servers on a schedule, unasked.


    Think of an application or server asset as a human being, and the services as the services needed to keep one healthy: dentistry, internal medicine, optometry, physical fitness.  A critical application is no different - it needs data center, OS management, middleware, devsecops and other services to be functioning well to remain secure.  One asset owner, potentially many services.


    Accountability sits with the asset owner.  Action responsibility sits with the service owner.  Sometimes those are the same group or person.  For mass services - OS patching across a thousand servers - they almost never are, and that is the normal shape of the work rather than a problem to be fixed.  The only defect is a blank - an asset with no owner, or a class of work on a group of assets which no service claims.


  3. Measure the estate risk, not the vulnerabilities.  (Paper 4 - Measuring What Matters.)  Estate Under Management is the share of the estate carried by a standing service.  Maturity is how reliably that service runs, on three binary properties: automated, no human involvement; preauthorized, it runs under a pre-agreed standard change model rather than per-instance negotiation; prescheduled, the window is fixed in advance.  All three earns Level 4, any two Level 3, any one Level 2, manual work Level 1.


    This system is illustrated in Figure 3 below.  Just visually, it can be seen where the biggest areas of opportunity lie - and it rolls up to a single estate score: each asset's risk times the maturity of the service carrying it, summed, over total asset risk times top maturity.  In simple and incorrect mathematical shorthand, Sum(asset risk x maturity) / (4 x Sum(asset risk)).  A fully automated estate scores one hundred percent, and unmanaged work drags the score down in proportion to the risk it carries.


    Approved exceptions are business choices, and the model treats them that way: they either leave the denominator entirely, or take a score of zero through four reflecting how well mitigated they are.  Two refinements belong to the deeper papers but deserve a line here: assets nest - applications contain servers, servers carry middleware - so ownership and scores roll up through sub-assets; and risk ownership is not remediation ownership.  The asset owner answers for the risk.  The services answer for the work.

  4. Exceptions are always choices.  (Paper 3 - Exceptions and Campaigns.)  Because every asset has an owner and every change required to maintain it belongs to a service (part 1), any part of the estate that is not being maintained appropriately has an accountable owner.  Where ITAM is immature, huge swaths may roll up to a CTO or CIO - which can make for good incentivization.


The important differences from traditional exceptions:

  • Scope is defined in ITAM terms, not vulnerability terms.  For example: Java SDKs in the plastics business, in states that start with a vowel, in towns that start with a consonant, on Red Hat servers - excluded from patching.  If ITAM cannot express the scope, the exception cannot exist.


  • Every exception is backed by a verified control - or by a named executive accepting the risk in writing.  A formally accepted risk stays in the model, scored zero through four on the mitigants standing behind it.  Risk nobody accepts rolls up to the CISO, CTO and/or CIO.


  • Mitigation is a state, not a clock.  Remediation is the clock, and it starts when a fix exists.  When disclosure arrives, the mitigating control is either already in front of the asset or it is not - a trade decided in advance, not during an incident.

 

SO WHY ISN'T EVERYONE DOING THIS?

The money is on the wrong side of the problem. Cybersecurity's budget buys scanners, analysts, dashboards and prioritization engines. Every dollar of it makes the report better.  Not one of them makes the estate safer. The teams who could make it safer are measured on uptime, delivery and cost - so saying no to you is the correct behavior on their part, and they are right.  Both sides spend rationally.  Neither spend closes the gap.


Economics named this decades ago - the split incentive, better known as the landlord-tenant problem: the landlord buys the furnace, the tenant pays the heating bill, and the efficient furnace never gets bought.5  With one difference that matters: a landlord and a tenant have no common boss.  A CIO and a CISO do - which is what makes this fixable, and what makes it inexcusable that it has not been fixed.


Inertia.  The current model is thirty years old, everybody is fluent in it, and nobody has ever been fired for running it.  A finding count can be produced on Monday morning.  Knowing what you actually own cannot.  The easier measure wins, the program stays busy enough to look healthy, and the harder question never gets asked.


Control, or the perception of it.  Cyber owns the scanner, so cyber is assumed to own the problem.  It is a myth and an expensive one.  The scanner is a measuring instrument, and owning the thermometer has never made anybody responsible for the weather.


Structures like this do not fix themselves.  They get fixed from above, and usually after something has already gone wrong.  Congress had to legislate the same defect in 1986: Goldwater-Nichols placed clear responsibility on the combatant commander and required that authority be "fully commensurate with the responsibility" for the mission.6  It took two failed operations to get there.


You do not have to wait for that.  And be clear about what the move below does not do.  It does not resolve the split.  It makes the split undeniable to the person who can.


So who moves first?  The CISO - and without giving anything up.  (Paper 6 - How To Get There.)  Cybersecurity keeps the accountability, the standard and the ledger.  The one thing that has to change is what it asks for.  Stop asking the CTO and the application teams how many findings they closed.  Require instead that every part of the estate be named to a service, with a scope, a committed time, and an honest answer on whether that time was met - then publish the share of the estate those services carry, with the excluded share beside it.  That demand costs no authority and needs no reorganization.  It can be made at the next steering meeting, and it puts the split into a number - the only thing that has ever moved a budget.  This is not a migration.  Nothing is replaced; the same work carries on with a name, an owner and a committed time attached.


DONE WELL, THIS IS A WIN-WIN-WIN-WIN

Nobody here is absorbing a security program.  Everyone on this list gets something they already wanted.


Make the case as the shared outcome it actually is.  Nobody has to lose for this one to work.

 

How many CISOs of today grew up wanting Beringer's authority with Lightman's skills, reigning over the big room at NORAD, but were ultimately greeted with death by PowerPoint, an endless supply of analytics, and a far different level of executive power - no scrambling of air force jets, launching ICBMs or scaring the shit out of everybody by declaring, "Sir, we are at DEFCON 1."


Sure, the occasional super code red could get the engines running for a tick, leading to really really insanely ridiculous conference calls with a hundred executives badgering two engineers for updates while they heroically work to control the issue.


Joshua said it best.  "The only winning move is not to play - that way."

TranSigma Consulting.  Richard Metz.  Method note available on request.

Start Delivering

Let's take action on your back-office problems

Corporate Headquarters

6 Corporate Dr. Suite 444

Shelton, CT 06484

  • LinkedIn

GET IN TOUCH WITH US

bottom of page