The model

The response half of threat-informed operations

TID-CMM asks one honest question — would we actually see it? TIR-CMM asks the question that follows: could we actually stop it? Not whether playbooks exist, but whether containment lands inside the adversary’s breakout window, with someone permitted to pull the trigger, at the blast radius intended — and whether any of it can be proved.

Could we actually stop it — inside the adversary’s breakout window, with someone permitted to pull the trigger, at the blast radius we intended — and can you prove it?

TIR-CMM is the response module of UTIOM and the companion to TID-CMM. It scores eight domains and 58 sub-capabilities, scores a Containment Lattice of eight attack-path stages crossed with eight asset classes, and applies seven integrity constraints mechanically — lower scores only, and not arguable.

From v0.2 it runs at three depths. Pulse is twenty questions in twenty minutes. Baseline is the 58 sub-capabilities, the lattice and response tempo. Assurance adds telemetry attributes, exercised scenarios and the governance layer. Each tier is a strict superset of the one below, so nothing on this page has to be answered before you are ready to answer it.

Three ways response fails

An organisation claiming 48.9% detection coverage can prove only 12.7%. That is the number TID-CMM produces. TIR-CMM exists because the question after it — could the organisation have acted on what it saw — has never been measured. Three failure modes, structurally parallel to TID-CMM’s three.

Playbooks without authority

The procedure exists. It is well written. At 03:00 on a Sunday, the analyst who could execute it is not permitted to isolate a production domain controller, and the person who is permitted is asleep and unreachable for forty minutes.

The playbook was never the constraint. Decision latency is the most under-measured variable in incident response, and it is invisible to every maturity model currently in use.

Response without attack paths

The IR plan is a generic NIST-shaped document: prepare, detect, contain, eradicate, recover, learn. It is not derived from this organisation’s crown jewels, modelled attack paths, or priority adversaries.

It tells you how to run an incident and nothing about which incidents you are structurally unable to stop. UTIOM’s Law 3 — crown jewels drive resource allocation — is violated the moment response is planned generically.

Capability without rehearsal

Playbooks are written, automation is built, and neither is ever executed against a clock. Tabletop exercises test the conversation, not the capability.

An unrehearsed playbook is an assumption wearing a document’s clothing. The failure surfaces exactly once, under maximum cost.

The consequence

Organisations invest in detection maturity while remaining structurally unable to act on what they detect. TID-CMM makes visibility honest. Without a matching instrument, response remains the last unmeasured link — and a validated detection that fires into an organisation that cannot contain is a very expensive alarm.

The two models sit on one axis: identical band arithmetic, no overlapping domains.
Model The question it asks What it produces
TID-CMM Would we actually see it? Validated Coverage Score — a claimed 48.9% against a provable 12.7%
TIR-CMM Could we actually stop it, in time, with authority, and can you prove it? Validated Response Score and Containment Margin — the engineered rate against the proven rate

Where it sits

UTIOM organises around three pillars. The two capability maturity modules map cleanly onto two of them, with no overlap.

Leadership & Governance

Vision · Strategy · Crown Jewels

UTIOM — maturity and capability assessments.

Engineering & Enablement

Threat Visibility · Threat Detection

TID-CMM — “would we see it?”

Operations & Analysis

Response · Continuous Improvement

TIR-CMM — “could we stop it?”

On the UTIOM V-model, TID-CMM measures the descending arm’s output — what the engineering decisions produced. TIR-CMM measures the ascending arm: whether the design survives contact and is proven by validation. This is why TIR-CMM’s validation domain carries the highest structural leverage in the model.

UTIOM Law 6 — operations functions as continuous incident response — is the load-bearing premise. TIR-CMM does not measure “the IR team.” It measures the organisation’s capacity to act, of which the IR team is one component and frequently not the binding one.

Division of scope

Question Answered by
Which adversaries matter to us? TID-CMM (TI) — consumed by TIR-CMM
What are our crown jewels and attack paths? UTIOM + TID-CMM (TM) — consumed by TIR-CMM
Would we observe the behaviour? TID-CMM (DC, DE, AV)
Does the alert reach a human or process correctly? Bridge — TID-CMM IR domain, redefined
Can we decide, contain, evict and restore? TIR-CMM
Are we faster than the adversary? TIR-CMM (Containment Margin)
Did we learn, and did it change detection? TIR-CMM (RG) → feeds back to TID-CMM

TIR-CMM deliberately contains no detection domain. It consumes detection maturity as an input constraint, R4 — Detection Dependency: RE ≤ D + 1, CE ≤ D + 1, FI ≤ D + 1, where D is the TID-CMM overall score if an assessment is imported, or an assessor-declared detection maturity flagged “unverified” otherwise. You cannot respond to what you never saw.

This is what makes the two models complementary rather than overlapping, and it is the design decision that keeps a combined assessment under two hours.

The eight domains

Weights total 100. Sub-capability counts total 58, matching TID-CMM’s structure so the two radars superimpose. Every sub-capability is scored 0–5. Sub-weights are within-domain and total 100 per domain.

All 58 are scored at Baseline and above. At Pulse, twenty of them are scored and the remaining domain scores are extrapolated from that subset.

ID Domain Weight Subs Core question
RP Response Preparation & Readiness 10% 7 Are we set up to run an incident at all?
RA Response Authority & Decision Rights 12% 6 Is someone permitted to act, in time?
RE Response Engineering & Playbooks 16% 10 Is response engineered, or written?
CE Containment, Eradication & Recovery 14% 8 Do we have graded options that work?
AO Automation & Orchestration 12% 7 Does machine speed reach the decision?
FI Forensics, Evidence & Investigation 10% 6 Do we know what actually happened?
RV Response Validation & Exercising 14% 7 Have we proven any of this?
RG Response Governance, Metrics & Improvement 12% 7 Does it improve, and can we report it?

RP Response Preparation & Readiness 10% · 7 subs

Are we set up to run an incident at all?

Whether the organisation can start an incident cleanly. Cheap to build, catastrophic to lack, and consistently over-scored.

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
RP-1 Incident response plan and scope 15 Plan is organisation-specific, names roles, covers cloud/identity/OT as applicable, reviewed annually Plan doc + version history + incident references
RP-2 Severity and classification model 15 Documented severity matrix driven by crown-jewel and business impact, applied consistently Severity matrix + sample of classified incidents
RP-3 Roles, rotas and 24×7 reachability 15 Named IR roles, documented rota, tested reachability, defined deputies Rota + unannounced call-out test records
RP-4 Response tooling readiness 15 Response tooling (EDR console, SOAR, forensic kit, out-of-band comms) is inventoried and access-tested Tool inventory + access test log
RP-5 Out-of-band communications and war-room 10 Documented out-of-band channel and bridge, contact tree maintained Exercise report showing OOB use
RP-6 Third-party, retainer and supplier readiness 15 IR retainer or in-house equivalent in place, SLAs known, onboarding pre-completed Retainer contract + activation/exercise record
RP-7 Legal, regulatory and communications preparation 15 Reporting obligations mapped (GDPR/DORA/NIS2/sector), legal counsel identified, holding statements drafted Obligation register + evidence of clock adherence

RA Response Authority & Decision Rights 12% · 6 subs

Is someone permitted to act, in time?

The domain no existing maturity model measures. Response fails here more often than it fails on tooling.

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
RA-1 Pre-authorised containment actions 25 A defined set of containment actions is pre-authorised by asset class and severity, documented and signed off by business owners Signed authorisation matrix mapped to lattice cells
RA-2 Decision latency measurement (MTTDecide) 20 MTTDecide is measured separately per severity and reported Metric series with timestamps from case system
RA-3 Escalation thresholds and triggers 15 Documented, objective escalation triggers tied to severity and crown-jewel involvement Trigger definitions + escalation audit trail
RA-4 Out-of-hours and degraded-mode authority 15 Named out-of-hours decision authority with documented deputies and a defined maximum response time Unannounced out-of-hours exercise record
RA-5 Business impact acceptance and risk ownership 15 Business owners have accepted, in advance and in writing, the disruption cost of defined containment tiers Signed impact-acceptance records
RA-6 Crisis and executive decision structure 10 Documented crisis management structure with defined activation criteria and executive decision rights Crisis activation log + executive exercise record

RE Response Engineering & Playbooks 16% · 10 subs

Is response engineered, or written?

The largest domain, mirroring TID-CMM's Detection Engineering weight. Response content deserves the same engineering discipline as detection content.

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
RE-1 Playbook coverage of the lattice 15 Every T1 and T2 lattice cell has an associated playbook or documented response procedure Playbook-to-cell coverage map
RE-2 Threat-informed derivation 12 Playbooks derive from modelled attack paths, crown jewels and priority actor TTPs Traceability matrix
RE-3 Playbook-as-code and version control 12 Playbooks are stored in version control with review, history and change approval Repository + CI pipeline evidence
RE-4 Playbook testing before release 12 Every playbook is executed in a test or staging context before production release CI test results
RE-5 Parameterisation and reusability 8 Playbooks are built from reusable, parameterised response actions rather than duplicated steps Action library + composition evidence
RE-6 Decision points and human gates 10 Playbooks contain explicit decision points, named decision owners, and defined default actions on timeout Playbook with instrumented decision points
RE-7 Playbook lifecycle and deprecation 8 Playbooks have owners, review cycles and a deprecation process Lifecycle register + health metrics
RE-8 Response action mapping to ontology 8 Response actions are mapped to RE&CT RA-codes and/or D3FEND techniques Mapping export
RE-9 Cross-domain response content (identity, cloud, OT, SaaS) 8 Playbooks exist for identity, cloud control plane and SaaS compromise, and OT where applicable Domain-specific playbooks + exercise evidence
RE-10 Response content sharing and reuse 7 Playbooks and actions are shared across teams/regions with a common standard Shared repository / contribution record

CE Containment, Eradication & Recovery 14% · 8 subs

Do we have graded options that work?

Graded, reversible options with known blast radius. Maps to D3FEND Isolate, Evict and Restore.

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
CE-1 Tiered containment options per asset class 16 Each asset class has graded containment tiers (observe / restrict / isolate / disable) with documented business impact Containment tier matrix + exercise data
CE-2 Blast radius control and reversibility 14 Blast radius is documented per action; every action has a defined rollback Rollback test records
CE-3 Identity containment (Tier-0) 14 Session revocation, credential reset, token invalidation and trust-break procedures exist and are owned Exercise record with timings
CE-4 Network and egress containment 10 Egress blocking, sinkholing and segment isolation are available and documented Automation + timing evidence
CE-5 Eradication completeness and re-entry prevention 14 Eradication addresses persistence, credentials and access paths, with a documented completeness checklist Eradication verification records + reinfection metric
CE-6 Recovery, restoration and integrity verification 12 Recovery procedures are integrated with IR, with defined RTO/RPO for crown jewels and integrity verification before return Recovery exercise report
CE-7 Backup resilience against destructive attack 10 Immutable or logically isolated backups exist for crown jewels, with separate credentials Restore test record
CE-8 Degraded-mode and business continuity operation 10 Documented degraded-mode operation for crown-jewel services, agreed with the business Continuity exercise record

AO Automation & Orchestration 12% · 7 subs

Does machine speed reach the decision?

Machine speed is only useful if it is permitted to reach the decision. Constrained by RA (R2).

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
AO-1 Enrichment and triage automation 14 Automated enrichment of alerts (asset, identity, intel, prior cases) before analyst contact Metric series
AO-2 Automated containment coverage 20 Automated containment exists for defined low-risk, high-confidence scenarios with pre-authorisation Automation coverage map + success rate
AO-3 Human-in-the-loop gate design 14 Gates are explicitly designed: which actions require a human, which do not, and what happens on timeout Gate design doc + latency data
AO-4 Integration coverage and API depth 14 Orchestration platform integrates with the tools controlling every in-scope asset class Integration inventory + health monitoring
AO-5 Automation reliability, safety and rollback 14 Automation has error handling, safety limits (rate/scope caps) and rollback paths Reliability metrics + safety test evidence
AO-6 Case management and workflow integrity 14 A case system holds all incidents with structured timeline, actions and timestamps Automated metric derivation from case data
AO-7 Automation change control 10 Automation changes go through review and testing before deployment CI/CD evidence

FI Forensics, Evidence & Investigation 10% · 6 subs

Do we know what actually happened?

Scoping accuracy determines eradication completeness. Under-invested almost universally.

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
FI-1 Triage and scoping capability 22 Structured triage produces a defensible scope: affected assets, identities and timeframe Scope-accuracy review records
FI-2 Evidence acquisition capability 18 Volatile and disk acquisition is possible for each in-scope asset class within a defined time Acquisition test records incl. cloud
FI-3 Chain of custody and evidence handling 15 Documented chain of custody suitable for internal disciplinary and insurance purposes Custody procedure + legal review
FI-4 Timeline reconstruction 15 Incident timelines are produced for significant incidents with source attribution Sample timelines + automation
FI-5 Log and evidence retention adequacy 15 Retention for crown-jewel-relevant sources exceeds median dwell time for priority actors Retention policy + successful lookback case
FI-6 Malware and artefact analysis access 15 Analysis capability available in-house or via retainer, with defined turnaround Analysis reports + resulting detection/playbook changes

RV Response Validation & Exercising 14% · 7 subs

Have we proven any of this?

The ceiling-setting domain. Everything else in the model is capped at RV + 1 (R1).

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
RV-1 Tabletop exercise programme 12 Regular tabletops covering crown-jewel scenarios with the right participants, findings tracked Exercise reports + remediation tracking
RV-2 Technical live-fire exercising 22 Playbooks are executed technically against emulated adversary activity in a controlled window Live-fire schedule + results per cell
RV-3 Purple team and end-to-end response validation 18 Purple team exercises continue through containment and eradication, not stopping at the alert End-to-end purple team reports
RV-4 Containment timing measurement under exercise 16 Exercises capture MTTD, MTTDecide and MTTC and compare them to breakout time Timing dataset across exercises
RV-5 Validation recency and coverage tracking 12 Validation status and date tracked per lattice cell; expiry applied Automated validation register
RV-6 Failure injection and resilience of the response function 10 Response is exercised under degraded conditions (key tool unavailable, key person absent, SIEM down) Degraded-mode exercise records
RV-7 Independent assessment and red team 10 Periodic independent red team or assessment includes response effectiveness, not just breach success Independent report + closure evidence

RG Response Governance, Metrics & Improvement 12% · 7 subs

Does it improve, and can we report it?

Whether the capability compounds. Closes UTIOM's Kaizen loop and feeds TID-CMM.

ID Sub-capability Wt What a 3 looks like Evidence for 4–5
RG-1 Response metrics programme 18 MTTD, MTTDecide, MTTC, MTTR and containment success are measured and reported Metric definitions + reporting history
RG-2 Post-incident review discipline 18 Blameless post-incident reviews occur for all significant incidents, with tracked actions PIR records + closure metrics
RG-3 Feedback loop into detection (TID-CMM) 16 Incident and exercise findings generate detection engineering work items Linked incident → detection change records
RG-4 Feedback loop into architecture and hardening 12 Findings generate hardening and architecture work items with owners Linked findings → architecture changes
RG-5 Response capability ownership and funding 12 Named owner and dedicated budget for response engineering, distinct from SOC staffing Budget + prioritisation evidence
RG-6 Regulatory and executive reporting 12 Defined reporting packs and regulatory notification procedures, with owners and clocks Reporting evidence with timestamps
RG-7 Continuous improvement cadence 12 A regular cadence reviews response capability against the model and adjusts the backlog Reassessment history

Evidence levels

v0.1 asked a yes/no question against every score: is there a named, dated artifact? v0.2 replaces that flag with a level of 0 to 3, and each level carries a cap on the score it will support. The levels and their caps are aligned to TID-CMM’s Validated Coverage grades, so an imported VC grade sets the evidence ceiling directly rather than having to be reinterpreted.

Defined in assets/model.js as EVIDENCE and EVIDENCE_CAP. Applied by constraint R3 at sub-capability level, before any domain score is computed.
Level VC grade Name What it expects Caps the score at What it means
0 VC0 Assertion only None, or a verbal claim. 1 The claim may be recorded, but readiness cannot be credited above 1.
1 VC1 Design or policy A document, policy, screenshot, or partial implementation. 2 Something exists on paper. Operational effectiveness is unproven.
2 VC2 Implemented and tested Implementation evidence plus a test result, ticket, or system record. 4 May reach Measured once tests and outcomes are current.
3 VC3 Repeatably validated Recent, repeatable, independently reviewable proof from the live environment. 5 May reach Adaptive where validation is sustained.

The cap is a ceiling, not a conversion

This is the point on which the whole mechanism turns, and it is the one most often misread. A high evidence grade permits a high score. It does not create one. VC3 evidence against a capability you score 2 leaves the score at 2 — excellent proof of a modest capability is still a modest capability. The cap only ever removes the part of a claim the evidence cannot carry.

Read the other way, the cap is why the evidence level is worth recording honestly. Raising a level to lift a cap does not improve the capability; it removes the finding. The gap between the score you would have given and the score the evidence permits is the finding.

A blank level is level 0

An evidence level left blank is treated conservatively as level 0 — assertion only, capping that sub-capability at 1. It is not waved through and it is not skipped. An assessment that declines to say what its evidence is has answered the question.

Records written against v0.1 still resolve: the old boolean flag meant “a named, dated artifact”, which is VC2, so a true flag imports as level 2 and caps at 4. Nothing has to be rescored.

Pulse does not ask for evidence levels at all. Twenty minutes is not enough time to review artifacts honestly, so the evidence cap is switched off for Pulse rather than applied against evidence nobody looked at. The reduced rigour is carried by the L3 band ceiling under R7 instead.

Attack-path stages

The rows of the Containment Lattice. Each carries a stage leverage λ, which weights a gap at that stage when the roadmap is ranked: a missing option early in the path costs more than a missing option late, because acting early is what changes the outcome.

v0.2 adds S0 — Prevent & Harden ahead of S1, and gives it the highest leverage in the model.

Defined in assets/model.js as STAGES. Full stage descriptions, the RRS scale and the expiry rules are on the lattice page.
Stage Name λ The question it asks
S0 Prevent & Harden New in v0.2 1.6 Can architecture stop it, or materially slow it, before response is needed?
S1 Initial Access & Foothold 1.5 Can we cut the entry before execution?
S2 Execution & Persistence 1.4 Can we kill it and prevent return?
S3 Privilege Escalation & Credential Access 1.3 Can we revoke and re-trust at speed?
S4 Discovery & Lateral Movement 1.2 Can we sever the path mid-flight?
S5 Collection & Staging 1.0 Can we interrupt before the data moves?
S6 Command & Control and Exfiltration 1.0 Can we block the egress and the channel?
S7 Impact & Objective 0.8 Can we limit, reverse and restore?

Prevention buys the one thing response cannot manufacture, which is time.

S0 is scored first, and it asks a question the rest of the lattice cannot: whether architecture or controls blunt the behaviour before a response is needed at all. It carries λ = 1.6, higher than any response stage, for a reason that is arithmetic rather than rhetorical. Every other stage improves the odds of winning a race that has already started. S0 shortens the race, or stops it being run. A model that scores only response will systematically undervalue the control that removed the incident, which is the opposite of the behaviour it should encourage.

It maps to ATT&CK Mitigations (M-codes) and D3FEND Harden, so a preventive control inventory already expressed in either vocabulary drops straight into the row. The lattice is therefore 8 stages × 8 asset classes = 64 cells at full extent in v0.2.

Three readiness lenses

The eight domains describe how the capability is built. That is the right structure for improving it and the wrong structure for reporting it: no executive has ever asked how the automation and orchestration domain is doing.

The lenses re-cut the same 58 sub-capabilities against the three outcomes leadership does ask about — can we respond, can we recover, can we keep running. No new questions are asked and no sub-capability is scored twice; this is a second view over the same scores, computed as a weighted mean of the R3-adjusted sub-capability values in each lens.

They are lenses, not partitions. A sub-capability can serve more than one outcome, and eight of them do: RP-3, RP-6, RA-4, RA-6, RE-9, CE-7, CE-8 and FI-1 each appear under two lenses. Nothing is double-counted in the overall score, which remains the domain-weighted figure; the lenses sit beside it, not inside it.

Lens Core question What it measures Subs
Readiness to respond If it started now, could we act on it in time? Authority to act, playbooks that exist and run, containment options with known blast radius, and the automation and investigation depth to use them. 27
Readiness to recover Could we get the business back, clean, and prove it? Eradication that actually removes the adversary, restoration with integrity verification, backups that survive a destructive attack, and the legal and communications machinery that runs alongside. 14
Operational resilience Could we keep running through it, and be harder to hit next time? Degraded-mode operation, a response function that survives losing its own tooling or people, and a feedback loop that converts every incident into a structural improvement. 15

What feeds each lens

Lens weights are 2 for a sub-capability the outcome depends on directly and 1 for one that contributes to it. They are lens weights only: they do not alter domain weights, sub-weights or the overall score. Forty-eight of the 58 sub-capabilities appear in at least one lens; the ten that do not — RE-5, RE-8, RE-10, AO-7, RV-2, RV-4, RV-5, RV-7, RG-1 and RG-5 — are engineering, validation and governance machinery that shows up in the domain scores and in the rehearsal ceiling rather than in an outcome.

Read down: every sub-capability feeding each lens, at its lens weight.
Lens Wt Sub-capabilities n
Readiness to respond 2 RA-1, RA-2, RA-4, RA-5, RE-1, CE-1, CE-2, CE-3, AO-2, FI-1, RP-3 11
1 RA-3, RA-6, RE-2, RE-3, RE-4, RE-6, RE-9, CE-4, AO-1, AO-3, AO-4, AO-6, FI-2, FI-4, RP-2, RP-4 16
Readiness to recover 2 CE-5, CE-6, CE-7, RP-6, RP-7 5
1 CE-8, FI-1, FI-3, FI-5, FI-6, RP-1, RG-6, RA-6, RE-9 9
Operational resilience 2 CE-7, CE-8, RV-6, RP-5, RA-4, RG-2, RG-3, RG-4 8
1 RV-1, RV-3, RP-3, RP-6, RG-7, RE-7, AO-5 7

Claimed and proven

Each lens reports twice. The claimed figure is the weighted mean as scored. The proven figure applies the rehearsal ceiling from constraint R1 — validation maturity plus one — because an outcome nobody has exercised is an outcome nobody can claim. A lens is not exempt from the constraint that governs every domain.

There is a consequence worth stating before anyone reads a lens chart. Where the rehearsal ceiling binds, it flattens all three proven figures to the same number, since they are all being cut by the same ceiling. Three equal readiness scores say nothing about the balance between responding, recovering and continuing to operate; they say that validation is the binding constraint on all three. In that state the claimed figures are what discriminate between the lenses — and the distance between claimed and proven is itself the finding, because it is the exact size of the readiness that has been engineered and never executed.

The lenses change the reporting, not the model. The band, the overall score, the seven constraints and the lattice are untouched by them. Their purpose is to let the same assessment answer a board question without translation, and to stop readiness being a word that means whatever the person saying it needs it to mean.

Maturity bands

Six bands, identical arithmetic ranges to TID-CMM so the two scores sit on one axis.

Band Range Name Definition
L0 0.00–0.99 Improvised Response is individual heroics. The outcome depends on who happens to be on shift. This is a build project, not an improvement project.
L1 1.00–1.99 Documented Plans exist on paper. Execution is unrehearsed and personality-dependent. The document has never met a clock.
L2 2.00–2.99 Repeatable Playbooks and tooling are in place and consistently executed for known scenarios, but remain unproven under time pressure. The most common band, and the band where response spending most exceeds response capability.
L3 3.00–3.99 Threat-informed Response traces to modelled attack paths and crown jewels. Containment is pre-authorised. Exercises are regular and evidenced. The realistic target for most organisations.
L4 4.00–4.99 Validated A standing validation function exists. Containment is demonstrably inside the breakout window for priority actors. Requires tempo evidence — an L4 claim without measured MTTDecide is not assessable.
L5 5.00 Adaptive Response adapts as the threat model changes; playbook and authority coverage regenerate from threat change without a project. Treat any L5 claim with scepticism unless the evidence is exceptional.

Band gates

Three gates sit above the arithmetic. They cap the band regardless of the decimal.

  • L4 and L5 are unreachable without supplied tempo evidence. Under R5 — Tempo Ceiling, a Tempo Ratio of 2.0 or above caps the band at L1; a ratio between 1.0 and 2.0 caps it at L2; tempo not supplied at all caps it at L3 and marks the report “tempo unverified.” A response capability structurally slower than the adversary is not a level 3 capability, whatever the documentation says.
  • L3 is unreachable with any Tier-1 blind cell. Under R6 — Blind-Cell Gate, a single in-scope cell at criticality T1 with a Response Readiness Status of 0 caps the band at L2. If there is a stage on a crown jewel where you have no response option at all, you are not threat-informed.
  • L5 is unreachable below the Assurance tier, and L5 is unreachable without independent review. Under R7 — Assessment Depth & Governance Ceiling, Pulse caps the band at L3 and Baseline caps it at L4; and a self-assessment run without evidence-led scoring, independent calibration and separation of assessor from approver caps at L4 whatever tier it was run at. Twenty questions cannot evidence adaptive maturity, and a result nobody challenged is not assurance.

The honest shock in TIR-CMM arrives in the band, not in the number. In the model’s numerical validation a self-assessed 2.69 became an adjusted 2.32 — a modest decimal move — while the reported band fell from L2 to L1 — Documented, because the gates fired.

Assurance tier

Telemetry attributes

Response depends on evidence, and evidence has properties beyond “we collect it.” Eight attributes are scored 0–5 against each asset class.

Defined in assets/tiers.js as TELEMETRY_ATTRS. Scored per asset class, not per data component.
Attribute The question it asks of an asset class
CollectionIs the evidence collected at all from this asset class?
CompletenessDoes it cover the whole estate, or only the parts that were easy?
TimelinessDoes it arrive fast enough to act on inside the breakout window?
Integrity & normalisationAre entities resolvable and timestamps trustworthy across sources?
RetentionDoes it survive longer than your median dwell time?
QueryabilityCan a responder actually search it under pressure, at speed?
Health monitoringWould you know if this source stopped, and how quickly?
Access protectionCan an adversary with a foothold alter or delete it?

Why per asset class, and not per data component

The conventional approach scores these attributes against each ATT&CK data component. That is 106 rows, and it is the single largest cost in a traditional assessment workbook. TIR-CMM asks the same eight questions against the eight asset classes instead: roughly eight rows rather than 106, for very little loss of signal.

The reason it loses so little is that telemetry health is a property of a platform, not of a log type. Retention, timeliness, health monitoring and access protection are set by the pipeline the platform feeds, so the answer is stable across every data component that platform emits. Scoring them 106 times mostly re-records the same eight answers at thirteen times the cost, and the variance it recovers is not what decides whether a responder can act.

The output is a per-asset-class telemetry mean and a named list of weak attributes — anything at 1 or below — which is what a remediation plan needs.

Assurance tier

Scenario validation

Questionnaires reveal claims. Scenarios reveal integration. A scenario traces one business-relevant path from first evidence to trusted recovery, and deliberately tests people, data, tooling, authority, suppliers and communications together. It is the only part of the model that can turn a claimed capability into a proven one, which is why an exercised scenario is the fastest route to lifting the rehearsal ceiling.

Twelve starter scenarios ship with the model, each scored across ten lifecycle stages.

Defined in assets/scenarios.js as SCENARIOS. Criticality 1–5; asset classes are the ones each scenario exercises.
ID Scenario Crit Asset classes
S01Enterprise ransomware and double extortion5A2, A3, A7
S02Identity, MFA and session-token compromise5A1, A4
S03Business email compromise4A1, A4
S04Cloud control-plane compromise5A4, A1
S05SaaS data theft and extortion5A4, A7
S06CI/CD and software supply-chain compromise5A6, A1
S07Edge device or VPN exploitation5A5, A3
S08Insider data removal4A7, A2
S09Destructive attack on backups5A7, A3
S10Managed provider or supplier compromise4A1, A3
S11OT or safety-critical disruption5A8, A5
S12Degraded-mode incident4A2, A3, A5

The ten lifecycle stages

Each in-scope scenario is scored 0–5 against every one of these. They are the seams a real incident tests, and they are where integration fails even when each component is individually healthy.

Defined in assets/scenarios.js as SCENARIO_STAGES.
Stage The question it asks
PlaybookDoes a usable, current procedure exist for this scenario?
Telemetry & detectionWould the behaviour produce evidence you would actually see?
Triage & investigationCan analysts confirm it, scope it, and explain what is affected?
Decision & authorityCan containment be authorised inside the response horizon?
ContainmentCan the behaviour be interrupted safely, at the intended blast radius?
Forensics & evidenceIs evidence preserved without delaying urgent risk reduction?
EradicationAre persistence, credentials and access paths actually removed?
RecoveryCan services, data, identities and control integrity be restored and trusted?
CommunicationsDo internal, executive, customer and regulatory comms run on time?
Third-party coordinationDo suppliers, retainers and partners perform to their agreements?

What makes a scenario pass

A scenario passes only when all four of the following hold. Three of the four are about the exercise rather than the answers, which is the point.

  • Every one of the ten stages scores above 1. Nothing at or below 1 anywhere in the lifecycle.
  • It was exercised, with a date.
  • The exercise was timed — an observed recovery time was recorded.
  • The observed recovery time met its target.

An exercise with no recorded times cannot lift a readiness claim. A scenario that was run but not measured is reported as run-but-unmeasured, and it counts as neither a pass nor a failure. Observed timings are what make a scenario proof rather than paperwork; without them the exercise records that people attended, not that the organisation can recover.

Pass conditions

The stage scores are judged against these seven conditions, which are what “good” means at each seam:

  • Required evidence is generated, retained, accessible and attributable to the correct entities.
  • Detection and case transitions occur within the approved response horizon.
  • Analysts can explain the hypothesis, affected scope, confidence, impact and next action.
  • Authorised containment is feasible, safe, logged, and reversible where it needs to be.
  • Forensic evidence is preserved without delaying urgent risk reduction beyond tolerance.
  • Recovery meets agreed objectives and validates identity, data, configuration and control integrity.
  • Every material deviation becomes a governed finding with an owner and a retest date.

The minimum scenario record

A scenario that cannot be handed to somebody else and re-run is a story. The minimum record is eight items:

  • Scenario ID, version, owner, business service, crown jewel and impact
  • Actor, access assumptions, platforms and explicit scope exclusions
  • Core ATT&CK techniques and environment-specific procedures
  • Required telemetry, analytics, case flow, entity joins and known blind spots
  • Investigation questions, decision thresholds and containment authority
  • Eradication criteria, clean-recovery requirements, RTO/RPO and integrity validation
  • Test method, injects, expected results, observed results, timestamps and evidence links
  • Findings, owner and retest date

Assurance tier

The governance layer

A populated assessment without controlled scope, evidence, decision rights and retesting is a draft, not an assurance result. This layer is what makes the difference, and it is the part maturity models most often omit.

Without evidence-led scoring, independent calibration and separation of assessor from approver, a result is a self-assessment. It is capped at L4 — useful, and often the right thing to run — but it must not be presented as assurance.

The eight assurance conditions

Answered yes or no. Three of them — GV-3, GV-4 and GV-5 — decide whether the result is assurance or self-assessment.

Defined in assets/governance.js as GOV_CHECKS. Governance completeness gates the confidence of the result, not its score.
ID Condition The question asked
GV-1Named executive sponsorIs there a named sponsor accountable for scope, risk acceptance and publication?
GV-2Scope frozen before scoringWas the scope and target agreed and frozen before any score was entered?
GV-3Evidence-led, not self-declaredWere scores set from evidence interviews and artifacts rather than self-declaration?
GV-4Independent calibrationHas a second reviewer calibrated every score above 3 and every critical gate?
GV-5Separation of assessor and approverIs the person proposing scores different from the person approving them?
GV-6Risk acceptances expireDoes every accepted risk carry an owner, a compensating control and an expiry date?
GV-7Closure requires retestIs closure evidence defined and independently retested before an action is closed?
GV-8Reassessment triggers definedAre the triggers that invalidate this assessment written down and monitored?

Decision rights

Who is entitled to do what to the assessment itself. The line that does most of the work is the third one.

Defined in assets/governance.js as DECISION_RIGHTS.
Role What they decide
Executive sponsor or delegated risk ownerApproves scope, target changes, risk acceptance, and publication of the final readiness statement.
Assessment leadCalibrates proposed scores, arbitrates disagreement, and owns the reconciliation log.
Technical and business ownersPropose scores and supply evidence. They do not approve their own high scores.
Independent validatorChallenges material claims, every score above 3, and every critical-gate closure.

The ten calibration questions

Asked of the assessor, not of the assessed. Each one is designed to find a score that is true on paper and false in practice.

  1. What would fail if the primary analyst, engineer, administrator or supplier contact were unavailable?
  2. Which score depends on a feature being licensed rather than implemented and tested?
  3. Which evidence proves the control works under the current architecture, and not only in a lab?
  4. What is the oldest evidence supporting a score above 3, and what has changed since it was collected?
  5. Which priority behaviour is visible but cannot be contained inside the required response horizon?
  6. Which recovery assumption has not been proven from an isolated or compromised identity state?
  7. Which cross-domain entity joins fail across identity, endpoint, cloud, SaaS, network and business evidence?
  8. Which alert or case cannot tell the analyst what is believed, why, what is affected, and what to do next?
  9. Which accepted risk lacks an expiry, a compensating control, or an accountable business owner?
  10. Which closed action has not been independently retested?

Response RACI

Decision rights during an incident, as distinct from decision rights over the assessment. This is the table the RA domain scores against, and the one organisations most often discover they do not have.

Defined in assets/governance.js as RACI. Role titles are indicative; the structure is not.
Decision Responsible Accountable Consulted Informed
Declare an incidentSOC shift leadIR leadService ownerCISO
Authorise endpoint containmentSOC analystIR leadService ownerIT operations
Authorise identity revocation (Tier-0)Identity platform leadCISOService owner, LegalExecutive sponsor
Authorise production service isolationIR leadBusiness service ownerCISO, IT operationsExecutive sponsor
Invoke recovery / continuityIT operationsBusiness service ownerIR leadExecutive sponsor
Regulatory notification decisionLegal & complianceExecutive sponsorCISO, DPOBoard
External communicationsCommunicationsExecutive sponsorLegal, CISOAll staff
Accept residual riskCISOBusiness service ownerEnterprise riskBoard
Close a critical gateAction ownerAssessment leadIndependent validatorExecutive sponsor

Anti-gaming rules

Six ways an assessment is quietly inflated, named so that they can be refused.

  • Do not count out-of-scope items in the denominator.
  • Do not score the availability of a product feature as an implemented capability.
  • Do not use stale tests, screenshots, policies or vendor statements as current operational proof.
  • Do not average away critical gates, failed scenarios, incomplete assessment or missing evidence.
  • Do not let the assessor be the only approver of their own high scores.
  • Do not raise an evidence level to lift a cap. The cap is the finding.

Governance rules

  • Critical-gate closure requires accepted evidence and a retest, not a status change in a tracker.
  • Risk acceptance must name a business owner, a rationale, a compensating control, an expiry and a reassessment trigger.
  • Risk acceptance expires. It never marks the underlying capability as complete.
  • A supplier statement may support evidence but never replaces customer-side proof of integration, access, containment and recovery.
  • The assessor must not also be the sole approver of high scores or of gap closure.
  • Freeze the imported baseline before scoring, so improvement is measured against a stable start.

Reporting to a board

  • Show scope, completion, evidence freshness, target and open critical gates beside every average.
  • Report business scenarios and survival outcomes before any coverage percentage.
  • Separate proven readiness, planned capability and accepted risk. Never merge them into one number.
  • Show a trend only when scope, method, target and data lineage are comparable; otherwise explain the break.
  • Use a small number of decision-oriented measures and link the detail to owners and actions.

Publication cautions

  • Do not publish stage-level weaknesses, evidence locations, rule identifiers, response authorities or recovery details outside their approved audience.
  • Remove sensitive system names and operational detail from external reports; keep the full evidence register in the controlled repository.
  • State the model version, assessment date, scope, exclusions and evidence limitations alongside any shared score.
  • TIR-CMM is a practitioner model. It is not certification, regulatory approval, or endorsement by MITRE, NIST, CISA or FIRST.