# TIR-CMM

## Threat-Informed Response Capability Maturity Model

**Model Specification v1.0**
**Date:** 24 August 2026
**Status:** Published
**Author:** Reza Adineh
**Position:** Response module of UTIOM (Unified Threat-Informed Operations Model). Companion and successor-in-scope to TID-CMM (Threat-Informed Detection Capability Maturity Model).
**Licence:** Model CC BY-ND 4.0 · Schemas and machine-readable model CC BY 4.0 · Tooling source-available, all rights reserved

### Artifact versions

TIR-CMM ships as four artifacts with independent version numbers:

| Artifact | What it is | Version at this release |
|---|---|---|
| **Model** | Domains, sub-capabilities, weights, lattice, constraints, bands | **1.0** — versioned by *this document* |
| **Browser tool** | The client-side assessment application and its scoring library | 1.0.0 |
| **Export schema** | The import/export contract other tools consume | 1.0 |
| **Workbook** | The offline spreadsheet that mirrors the tool's arithmetic | versioned separately, and states the model version it implements |

**This document versions the model only.** The four numbers are deliberately allowed to diverge. They change for different reasons and at different rates: a tool release that fixes a rendering defect must not imply the model moved, and a model change that adds a constraint must not force a schema bump on consumers whose parsing is unaffected. Forcing them to match would either freeze one artifact or falsely announce a change in the others.

They are therefore never presented as one number. Every export states all of them separately, so a result can always be read against the exact model that produced it, the exact tool that computed it and the exact contract it was written to. Do not infer that a model version and a tool version sharing a digit are related, and do not treat a mismatch as an error.

**The machine-readable model is canonical.** Where this document and the published machine-readable model disagree on a domain name, a sub-capability name, a weight, a level-3 anchor, an evidence descriptor or a constraint threshold, the machine-readable model is correct and this document is defective. The engine is what an assessment actually runs against; prose is a description of it.

---

## 0. Reader's map

| Section | What it settles |
|---|---|
| 1 | Why the model exists — the three response failure modes |
| 2 | Where it sits inside UTIOM and how it bonds to TID-CMM |
| 3 | The Containment Lattice — the atomic scoring unit, and what it does not measure |
| 4 | The two headline metrics: VRS and Containment Margin, and how tempo is defined |
| 8.4.3 | The prerequisite check — running TIR-CMM without TID-CMM |
| 5 | Maturity bands L0–L5 |
| 6 | The eight domains and 58 sub-capabilities |
| 7 | Scoring mathematics |
| 8 | Integrity constraints R1–R7, including R5b |
| 8A | Readiness lenses — respond, recover, withstand |
| 9 | Prioritisation, the blueprint and roadmap generation |
| 10 | Standards crosswalk |
| 11 | The TID-CMM bridge — resolving the IR domain overlap |
| 12 | Assessment flow (ten steps) |
| 13 | Data model and interop contract |
| 14 | Consistency checks — informative, never normative |
| 15 | Known limitations and open questions |
| 16 | Release status |
| Annex A | OT/ICS profile — normative extension, **not yet validated** |

---

## 0A. What this model is not

Stated once, at the front, because every subsequent section depends on it.

- **Not a certification.** There is no accrediting body, no register of assessed organisations, no pass mark and no badge. A TIR-CMM result is a self-produced instrument reading, and its own R7 constraint records how far it can be trusted.
- **Not an industry benchmark.** The model publishes no distribution of scores, because none exists. Any statement of the form "the average organisation scores X" is not sourced from this model.
- **Not a risk-quantification model.** It produces no loss estimate, no probability and no monetary figure. Containment Margin is a time measurement, not a risk metric.
- **Not a replacement for UTIOM.** TIR-CMM is one module of UTIOM and measures one pillar of it. It does not assess vision, strategy, crown-jewel definition or threat intelligence.
- **Not a product-selection framework.** No sub-capability names a vendor, a product category or a tool. The roadmap output names capabilities and authorities required, never products.

And, equally, what it is:

- **Usable standalone.** TIR-CMM does not require TID-CMM. The prerequisite check (§8.4.1) supplies `D` for constraint R4, at the cost of a band ceiling.
- **Compatible with TID-CMM.** Where a TID-CMM assessment exists, it imports directly and the two scores sit on one axis.
- **Browser-local.** The tool is entirely client-side. No server calls, no analytics, no transmission of assessment data. This is a design constraint, not a feature, and it is why the model can be run against real estate data.

---

## 0B. What changed in v1.0

v1.0 closes the gap between what v0.2 said the model did and what the engine actually did. Four of the six changes below are implementations of rules v0.2 had already published as normative and never enforced. That is stated plainly rather than presented as new capability: an unenforced rule is a defect, and correcting it is the release.

| Change | Section | Nature |
|---|---|---|
| **R5b — estimated tempo caps at L3** | §8.5.1 | Published as model behaviour in v0.2; **unimplemented until v1.0**. An archetype breakout of 1440 minutes previously produced a Tempo Ratio of 0.03 and permitted L5 on a guess. |
| **§3.6 recency is enforced** | §3.6 | Published as normative in v0.2; **unimplemented until v1.0**. A cell marked Proven stayed Proven for ever and VC3 evidence never aged. Both now revert. |
| **Tempo intervals defined by start and stop events** | §4.2.1 | New. Two assessors measuring different things produced incomparable results; the intervals are now defined by the events that open and close them. |
| **Breakout source hierarchy** | §4.2.3 | New. Five ranks, of which two count as measured. This is what R5b tests. |
| **R4 basis settled** | §8.4 | Open question in v0.2, now closed: `D` is TID-CMM's constraint-adjusted, pre-substitution **overall** score, and the telemetry domain is not a second cap. |
| **S0 scope boundary defined** | §3.1.1 | New. S0 was added in v0.2 without a boundary, and without one it drifts into a general preventive-control assessment. |

Two further additions are documentation of behaviour that already shipped and was justified only in code comments: the renormalisation of domain weights when a domain is unanswered, and Pulse's bypass of the R3 evidence cap (§7.2). Two are new material: the consistency-check layer (§14) and the OT/ICS profile (Annex A), the latter explicitly **not yet validated**.

### The migration consequence, stated plainly

An assessment carried over from v0.2 has no proof dates, because v0.2 never asked for them. Under v1.0 that assessment's Proven lattice cells revert to Engineered and its VC3 evidence grades down to VC2, until dates are added.

This is not a regression and it is not a scoring change made to depress results. It is the intended behaviour of an honesty mechanism that was normative but unenforced. A cell whose last proof cannot be dated cannot be shown to be current, and "we proved this at some point" is precisely the claim §3.6 exists to refuse. Re-dating an assessment is a data-entry exercise of minutes; the alternative is a model that credits validation nobody can locate.

---

## 0C. What v0.2 established, and remains in force

v0.2 responded to two things: a review of the UTIOM-ORR blueprint and its companion workbook, and the plain observation that v0.1 was too heavy to start with. Its structural decisions are unchanged in v1.0 and are restated here because the rest of the document depends on them.

**Tiering, not simplification.** One model, three depths, each a strict superset of the one below, so nothing is ever rescored:

| Tier | Time | Covers | Band ceiling |
|---|---|---|---|
| **Pulse** | 20 minutes | 20 questions, no lattice, no evidence levels | **L3** |
| **Baseline** | 1–2 hours | All 58 sub-capabilities with evidence levels, the lattice, authority map, tempo | **L4** |
| **Assurance** | 2–4 weeks | Adds telemetry attributes, exercised scenarios, the evidence register and the governance layer | none |

The ceiling is not a punishment. **The depth an assessment reaches is itself evidence** about how far its result can be trusted. Twenty questions cannot evidence a claim of adaptive maturity, however good the answers are.

| Established in v0.2 | Detail |
|---|---|
| **Evidence 0–3 replaces the binary flag** | VC0 caps at 1, VC1 at 2, VC2 at 4, VC3 at 5 — aligned to TID-CMM's Validated Coverage so an imported grade sets the ceiling directly. The cap is a **ceiling, not a conversion**: a high evidence grade permits a high score, it does not create one. A blank level is treated conservatively as VC0. |
| **S0 Prevent added to the lattice** | Ahead of S1, carrying the highest stage leverage (λ 1.6), because prevention buys the one thing response cannot manufacture, which is time. The lattice is 8 × 8 = 64 maximum cells. Its scope boundary is defined in v1.0 at §3.1.1. |
| **Scenario validation protocol** | 12 starter scenarios, 10 lifecycle stages each, pass conditions, minimum scenario record. A scenario passes only when every stage is above 1, it was exercised, the exercise was **timed**, and observed recovery met its target. An untimed exercise cannot lift a readiness claim. |
| **Telemetry attributes** | The eight attributes — collection, completeness, timeliness, integrity, retention, queryability, health, protection — scored **per asset class** rather than per ATT&CK data component. Telemetry health is a property of a platform, not of a log type. |
| **Governance layer** | Eight assurance conditions, the ten calibration questions, decision rights, RACI, anti-gaming rules, board reporting rules and publication cautions. |
| **R7 — assessment depth and governance ceiling** | Pulse caps at L3; Baseline at L4; and at Assurance depth, a result without evidence-led scoring, independent calibration and separation of assessor from approver is a self-assessment and caps at L4. |

**Where TIR-CMM diverges from the blueprint, deliberately:**

1. **Detect is not weighted at 20%.** The blueprint gives Detect the highest of its eight lifecycle-stage weights, in a module that by design sits downstream of a detection assurance model. That double-counts TID-CMM and lets a strong detection score inflate a response result — precisely the lifecycle asymmetry the module exists to expose. TIR-CMM has no detection domain at all; detection enters solely as constraint R4.

2. **Critical gates stay rare.** The workbook flags 102 of 180 questions as critical gates. When 57% of questions can independently fail an assessment, the gate stops being a signal. TIR-CMM's band gates are R5 (with R5b), R6 and R7 — each with a distinct and defensible trigger.

3. **Validation is not weighted at 5%.** The blueprint opens its scenario section with "questionnaires reveal claims; scenarios reveal integration" and then weights Validate lowest of the eight. TIR-CMM inverts that: RV carries 14% and, through R1, sets the ceiling for every other domain.

---

## 1. The problem TIR-CMM exists to solve

TID-CMM asks one honest question — *"would we actually see it?"* — and produces one uncomfortable number: an organisation claiming 48.9% detection coverage can prove only 12.7%.

TIR-CMM asks the question that follows, and it is not *"do you have playbooks."* It is:

> **Could we actually stop it — inside the adversary's breakout window, with someone permitted to pull the trigger, at the blast radius we intended — and can you prove it?**

Three failure modes, structurally parallel to TID-CMM's three:

| TID-CMM's failure mode | TIR-CMM's parallel |
|---|---|
| **Coverage without visibility** — rules exist and map to techniques, but the data source is absent or degraded | **Playbooks without authority** — the procedure exists and is well written, but nobody may execute it at 03:00 |
| **Detection without threat modelling** — content comes from generic subscriptions, not an organisation-specific adversary profile | **Response without attack paths** — a generic IR plan, not derived from this organisation's crown jewels and modelled paths |
| **Capability without evidence** — detection assumptions remain untested until a breach tests them | **Capability without rehearsal** — playbooks and automation are never executed against a clock |

### 1.1 Playbooks without authority

The procedure exists. It is well written. At 03:00 on a Sunday, the analyst who could execute it is not permitted to isolate a production domain controller, and the person who is permitted is asleep and unreachable for forty minutes. The playbook was never the constraint. **Decision latency is the most under-measured variable in incident response**, and it is invisible to every maturity model currently in use.

### 1.2 Response without attack paths

The IR plan is a generic NIST-shaped document: prepare, detect, contain, eradicate, recover, learn. It is not derived from this organisation's crown jewels, this organisation's modelled attack paths, or this organisation's priority adversaries. It tells you *how to run an incident* and nothing about *which incidents you are structurally unable to stop*. UTIOM's Law 3 — crown jewels drive resource allocation — is violated the moment response is planned generically.

### 1.3 Capability without rehearsal

Playbooks are written, automation is built, and neither is ever executed against a clock. Tabletop exercises test the conversation, not the capability. An unrehearsed playbook is an assumption wearing a document's clothing. The failure surfaces exactly once, under maximum cost.

### 1.4 The consequence

Organisations invest in detection maturity while remaining structurally unable to act on what they detect. TID-CMM makes visibility honest. Without a matching instrument, response remains the last unmeasured link — and a validated detection that fires into an organisation that cannot contain is a very expensive alarm.

---

## 2. Position within UTIOM

UTIOM organises around three pillars. The two capability maturity modules map cleanly onto two of them, with no overlap:

```
┌─────────────────────────────────────────────────────────────────┐
│  UTIOM — Unified Threat-Informed Operations Model               │
├──────────────────┬────────────────────┬─────────────────────────┤
│  LEADERSHIP &    │  ENGINEERING &     │  OPERATIONS &           │
│  GOVERNANCE      │  ENABLEMENT        │  ANALYSIS               │
│                  │                    │                         │
│  Vision          │  Threat Visibility │  Response               │
│  Strategy        │  Threat Detection  │  Continuous Improvement │
│  Crown Jewels    │                    │                         │
│                  │                    │                         │
│  ── UTIOM ──     │  ── TID-CMM ──     │  ── TIR-CMM ──          │
│  maturity +      │  "would we         │  "could we              │
│  capability      │   see it?"         │   stop it?"             │
│  assessments     │                    │                         │
└──────────────────┴────────────────────┴─────────────────────────┘
```

On the UTIOM V-model, TID-CMM measures the descending arm's output (what the engineering decisions produced); TIR-CMM measures the ascending arm (whether the design survives contact and is proven by validation). This is why TIR-CMM's validation domain carries the highest structural leverage in the model.

**UTIOM Law 6 — operations functions as continuous incident response** — is the load-bearing premise. TIR-CMM does not measure "the IR team." It measures the organisation's capacity to act, of which the IR team is one component and frequently not the binding one.

### 2.1 Division of scope

| Question | Answered by |
|---|---|
| Which adversaries matter to us? | TID-CMM (TI) — consumed by TIR-CMM |
| What are our crown jewels and attack paths? | UTIOM + TID-CMM (TM) — consumed by TIR-CMM |
| Would we observe the behaviour? | TID-CMM (DC, DE, AV) |
| Does the alert reach a human or process correctly? | **Bridge** — TID-CMM IR domain, redefined (§11) |
| Can we decide, contain, evict and restore? | **TIR-CMM** |
| Are we faster than the adversary? | **TIR-CMM** (Containment Margin) |
| Did we learn, and did it change detection? | TIR-CMM (RG) → feeds back to TID-CMM |

TIR-CMM deliberately contains **no detection domain**. It consumes detection maturity as an input constraint (R4). This is what makes the two models complementary rather than overlapping, and it is the design decision that keeps a combined assessment under two hours.

---

## 3. The Containment Lattice — the atomic scoring unit

TID-CMM scores coverage against in-scope ATT&CK techniques. TIR-CMM does not, and the reason is structural rather than a simplification.

**Response readiness is primarily assessed by attack-path stage and asset class, because containment primitives are shared across techniques on the same asset.** Revoking a session, isolating an endpoint, freezing a pipeline or breaking a trust are the same engineered actions, executed by the same tooling, authorised by the same person, whichever technique produced the need for them. What varies is *where on the attack path you are* and *what kind of asset you are acting on* — and those are the two axes the model scores.

The atomic unit of TIR-CMM is therefore a **lattice cell**: one attack-path stage crossed with one asset class.

This is a statement about the scoring dimensions, not a claim that response never varies by technique. It plainly does at the procedural level — evicting a kernel-mode driver and evicting a malicious OAuth grant are different work — and §3.7 defines where that detail is recorded.

### 3.1 Attack-path stages (rows)

Eight stages, collapsed from ATT&CK tactics into the boundaries at which a *different response decision* is available:

| ID | Stage | Response question | ATT&CK tactics folded in |
|---|---|---|---|
| **S0** | Prevent & Harden | Can architecture stop it, or materially slow it, before response is needed? | ATT&CK Mitigations (M-codes); D3FEND Harden |
| **S1** | Initial Access & Foothold | Can we cut the entry before execution? | Initial Access, Resource Development |
| **S2** | Execution & Persistence | Can we kill it and prevent return? | Execution, Persistence, Defense Evasion |
| **S3** | Privilege Escalation & Credential Access | Can we revoke and re-trust at speed? | Privilege Escalation, Credential Access |
| **S4** | Discovery & Lateral Movement | Can we sever the path mid-flight? | Discovery, Lateral Movement |
| **S5** | Collection & Staging | Can we interrupt before the data moves? | Collection |
| **S6** | Command & Control and Exfiltration | Can we block the egress and the channel? | Command and Control, Exfiltration |
| **S7** | Impact & Objective | Can we limit, reverse and restore? | Impact |

**Stage leverage** (used in prioritisation, §9): containing early is worth more than containing late.

| Stage | S0 | S1 | S2 | S3 | S4 | S5 | S6 | S7 |
|---|---|---|---|---|---|---|---|---|
| Leverage λ | 1.6 | 1.5 | 1.4 | 1.3 | 1.2 | 1.0 | 1.0 | 0.8 |

S0 carries the highest leverage: an attack path that architecture closes never needs a response at all.

S7 is weighted lowest not because impact response is unimportant, but because by S7 the model is measuring damage limitation rather than defence. An organisation whose entire response capability is concentrated at S7 is running a disaster recovery function and calling it incident response — the leverage weighting surfaces this automatically.

#### 3.1.1 The S0 boundary

S0 has the widest possible failure mode: left undefined, it swallows the whole of preventive security and turns a response model into a control audit. The boundary is therefore stated as a test, not as a topic list.

> **S0 Prevent & Harden measures pre-incident attack-path interruption that materially changes response readiness or buys response time.**

A control belongs in S0 only if you can answer this question concretely: *when this control holds, which response action does the organisation no longer need to take, or how much longer does it have to take it?* If the answer is "the risk is lower", the control is real and it is out of scope here.

**S0 is not a general preventive-control assessment, not a vulnerability-management assessment and not an exposure-management assessment.** Those are legitimate disciplines with their own instruments; UTIOM assesses them elsewhere. Scoring them here would double-count them and would let a strong hygiene programme inflate a response result — the same asymmetry §2.1 keeps detection out for.

**In scope — architecture that changes the response problem:**

| Example | Why it is S0 |
|---|---|
| Tiered administration that prevents a workstation credential from reaching Tier 0 | Removes the S3 → A1 escalation path entirely. The response action that would have been needed no longer exists. |
| Egress control that removes a viable C2 channel | Deletes an S6 containment race. There is no channel to block under pressure because there was never a channel. |
| Immutable or logically isolated backups for crown jewels | Removes the recovery decision. Encryption stops being a negotiation and becomes a restore. |
| Network segmentation that confines lateral movement to a single zone | Bounds blast radius before an incident, so S4 containment has a smaller problem and more time. |
| Phishing-resistant authentication on Tier-0 and crown-jewel access paths | Removes credential replay as a foothold mechanism, which removes the S1 → S3 sequence that most response time is spent on. |

**Out of scope — preventive security that does not change the response problem:**

| Example | Why it is not S0 |
|---|---|
| Patch cadence and mean time to patch | A hygiene and exposure metric. It changes the probability that a path is available, not the response actions available on it. |
| Open CVE counts and severity distribution | A vulnerability-management output. It names no attack-path stage, no asset-class containment primitive and no response time saved. |
| Phishing-simulation click rates | An awareness metric about people's behaviour, not an architectural interruption of a path. |
| General hardening baselines and benchmark compliance | Measures conformance to a standard configuration. Only the specific baseline items that close a modelled path qualify, and they qualify individually, not as a percentage. |
| Security-awareness training completion | Measures delivery of training, not the interruption of any modelled path. |

Where a control is genuinely borderline, the tie-break is the recording requirement: an S0 cell must be able to name the modelled attack path it interrupts. A control that interrupts no modelled path is not an S0 cell in this assessment, whatever its merit.

### 3.2 Asset classes (columns)

Eight classes, derived from the "what you run" and "what you protect" steps. Classes not present in the environment are excluded entirely.

| ID | Asset class | Typical containment primitives |
|---|---|---|
| **A1** | Identity & Access Infrastructure (Tier-0: AD, Entra ID, IdP, PAM, federation) | Disable, revoke session/token, force re-auth, break trust, tier-isolate |
| **A2** | Endpoint & User Compute (laptops, desktops, VDI, mobile) | EDR network-isolate, kill process, quarantine, reimage |
| **A3** | Server & Datacentre Workload (on-prem and IaaS servers, hypervisors, containers) | Isolate, snapshot, quiesce, failover, rebuild |
| **A4** | Cloud Control Plane & SaaS (cloud management planes, SaaS tenants, OAuth apps) | Revoke key/role, quarantine principal, disable API, tenant restrict |
| **A5** | Network & Edge (VPN, firewall, proxy, DNS, email gateway) | Block, sinkhole, ACL, disable tunnel, quarantine mail |
| **A6** | Application, Source & CI/CD (source repos, build pipelines, artefact stores, signing) | Freeze pipeline, revoke signing key, roll back release, disable webhook |
| **A7** | Data Stores & Backup (databases, file estates, object stores, backup systems) | Restrict access, immutable snapshot, restore, integrity verify |
| **A8** | OT / ICS / IoT & Specialist (industrial control, building systems, medical, specialist) | Segment, safe-state, manual fallback |

A8 carries OT in the **core** lattice, and it is a full column with the same scoring rules as every other class. The optional OT/ICS profile in Annex A adds depth to that column; it does not create the coverage.

### 3.3 Lattice sizing and scoping

Maximum lattice: 8 × 8 = **64 cells**. This is deliberately never the working set.

A cell is **in scope** only when *both* conditions hold:

1. The asset class exists in the environment (declared in step 1), **and**
2. At least one modelled attack path — imported from TID-CMM or declared locally — traverses that stage on that asset class.

Typical in-scope lattice: **24–38 cells**. A cell outside scope is marked N/A and excluded from both numerator and denominator, exactly as TID-CMM handles N/A sub-capabilities — no penalty, no benefit.

This is the mechanism that keeps the assessment inside the one-hour promise while giving the model far richer output than a flat questionnaire.

### 3.4 Cell criticality tiers

Each in-scope cell inherits a tier from the crown-jewel mapping:

| Tier | Meaning | Weight ω |
|---|---|---|
| **T1** | The cell *is* a crown jewel, or directly hosts one | 3 |
| **T2** | The cell lies on a modelled attack path to a crown jewel | 2 |
| **T3** | In scope, but peripheral to crown-jewel paths | 1 |

### 3.5 Response Readiness Status (RRS) — 0 to 3 per cell

Deliberately identical in shape to TID-CMM's validated-coverage status, so the two lattices read the same way.

| Status | Name | Definition |
|---|---|---|
| **0** | **No option** | No means exists to act at this stage on this asset class. You would improvise or watch. |
| **1** | **Manual only** | An action is technically possible but ad-hoc, undocumented, dependent on a specific individual, or gated behind an approval path with no defined SLA. |
| **2** | **Engineered** | A documented, parameterised playbook exists; tooling can execute it; an owner is named; authority is defined in advance. **Unproven.** |
| **3** | **Proven** | Executed against a real incident or live-fire exercise within the recency window, met its stage time objective, and the blast radius was as designed. |

### 3.6 Status 3 expires — and expires faster than detection

TID-CMM expires validation at eighteen months. **TIR-CMM expires at twelve months**, because response capability rests on people and authority, and organisations change those faster than they change detection logic.

**This rule was normative in v0.2 and was not implemented until v1.0.** Until this release a cell marked Proven stayed Proven for ever, and VC3 evidence never aged. Both now revert.

#### 3.6.1 Lattice cells

A cell at status 3 reverts to status 2 on **any** of the following:

1. The recorded proof date is more than **twelve months** before the assessment date.
2. **No proof date is recorded at all.** An undated proof cannot be shown to be current, so it is treated as expired rather than assumed valid.
3. Any of the five declared structural triggers applies. These are declared by the assessor, per cell:

| Trigger | Declaration |
|---|---|
| `tooling` | The tooling that executes the action changed (EDR, SOAR, IdP) |
| `oncall` | The on-call or escalation model changed |
| `reorg` | The response function was reorganised, or the named owner left |
| `architecture` | A material architecture change affected this asset class |
| `provider` | The managed provider or retainer delivering the action changed |

Recency only ever moves a cell **down**, and only ever from 3 to 2. Nothing in this rule can raise a status.

The declared status is retained alongside the effective one, so a report shows what expired rather than silently reporting a lower number. The reported proven rate is computed on effective status; the declared proven rate is reported next to it, and the gap between them is a validation-debt figure in its own right.

#### 3.6.2 Evidence levels

The same principle applies to evidence, because response evidence perishes for the same reason. **VC3 evidence grades down to VC2** — with its cap falling from 5 to 4 — where the evidence date is missing or more than twelve months old.

VC3 means *recent, repeatable, independently reviewable proof*. Undated proof fails the first word of that definition. Grading to VC2, "implemented and tested", is the accurate description of what an aged proof still demonstrates.

#### 3.6.3 Migration from v0.2

v0.2 never asked for dates, so a v0.2 assessment carried into v1.0 has none. Its Proven cells will therefore revert to Engineered and its VC3 evidence will grade to VC2 on first computation under v1.0.

**This is the intended behaviour**, not a defect and not a recalibration. The remedy is to record the dates; where a date cannot be found, the honest reading is that the proof cannot be produced, which is exactly what the lower status now says. Response maturity is perishable in a way detection maturity is not, and this rule is the single most important honesty mechanism in the model.

### 3.7 ATT&CK and ontology traceability beneath a cell

Response readiness is scored on stage and asset class. Everything more granular is **optional traceability recorded beneath a lattice cell**:

| Reference | Typical use beneath a cell |
|---|---|
| ATT&CK techniques and sub-techniques | Which observed behaviours this cell's response has actually been exercised against |
| Procedures | The specific implementation a priority actor uses, where it changes the containment step |
| Actors | Which prioritised adversary drove this cell into scope |
| ATT&CK Mitigations (M-codes) | Which mitigations the S0 row's architectural interruptions correspond to |
| D3FEND techniques | The countermeasure vocabulary for the containment primitive used |
| Detection use cases | The TID-CMM content that would raise the alert this cell responds to |

These references are **never mandatory scoring dimensions and never alter a cell's status**. A cell with forty mapped techniques and a cell with none score identically if their response readiness is identical. The references exist for three purposes only: deriving lattice scope from TID-CMM's technique scope, evidencing that an exercise covered what it claimed, and explaining a score to somebody who thinks in ATT&CK.

An assessment that records none of them is complete and valid. An assessment that records all of them has better provenance and the same score.

---

## 4. The two headline metrics

### 4.1 Validated Response Score (VRS)

The direct analogue of TID-CMM's Validated Coverage Score. Computed on **effective** RRS, after §3.6 recency has been applied.

**Unweighted:**

```
VRS = Σ RRS(c) / (3 × N)          for all in-scope cells c, N = |in-scope cells|
```

**Crown-jewel weighted (the reported figure):**

```
VRS_cj = Σ (ω(c) × RRS(c)) / (3 × Σ ω(c))
```

Reported alongside it, and expected to be the number that lands in the room:

- **Engineered rate** = proportion of cells at status ≥ 2 — "we believe we can act here"
- **Proven rate** = proportion of cells at effective status 3 — "we have shown we can act here, recently"
- **Declared proven rate** = the same figure before recency, reported next to it — the difference is validation debt
- **Blind cells** = count of status-0 cells at tier T1 or T2 — *the places on a path to a crown jewel where you have no option at all*

The expected finding, mirroring TID-CMM's 48.9% / 12.7%, is a large gap between engineered rate and proven rate. The model is built to make that gap unavoidable rather than to flatter it.

### 4.2 Containment Margin (CM) — the metric unique to response

Detection maturity is a coverage problem. Response maturity is a **race**. A model that scores response without reference to adversary tempo is measuring paperwork.

For each priority actor *a* with breakout time *B(a)*:

```
CM(a) = B(a) − ( MTTD + MTTDecide + MTTC )
```

```
TR(a) = ( MTTD + MTTDecide + MTTC ) / B(a)          "Tempo Ratio"
```

**MTTDecide is the contribution.** Every maturity model in circulation folds decision latency into MTTR, where it disappears. Separating it exposes the most common and most fixable failure in incident response: the organisation is not slow at containing, it is slow at being *allowed* to contain. Organisations that measure it routinely discover MTTDecide exceeds MTTC by an order of magnitude.

#### 4.2.1 The intervals, defined by their start and stop events

Two assessors must measure the same thing or nothing is comparable, and "mean time to detect" is measured at least four different ways in common practice. The following definitions are **normative**. The tool and this specification render them from the same table in the machine-readable model.

| Interval | Starts when | Stops when |
|---|---|---|
| **MTTDetect** | The adversary action occurs in the environment. | A human or system has validated the activity as a real incident requiring a response decision. |
| **MTTDecide** | The activity has been validated as requiring a containment decision. | The authorised containment decision is recorded by the person entitled to make it. |
| **MTTContain** | The authorised containment decision is recorded. | The containment action has taken effect and the adversary has lost the capability in question. |
| **Breakout time** | The adversary establishes an initial foothold. | The adversary moves laterally to a second system. |

Three consequences follow from those boundaries, and each one moves a real organisation's number:

- **MTTDetect is not the time an alert fired.** An unread alert has not detected anything. The clock stops at validation, not at generation.
- **MTTDecide is the interval this model exists to expose.** It disappears inside MTTR, and it is where most of the loss sits.
- **MTTContain ends when the action works, not when it was issued.** A disable that takes twenty minutes to propagate is twenty minutes.

Breakout time is the **adversary's** clock, not yours. In most organisations it is sourced rather than measured, which is what §4.2.3 and constraint R5b govern.

#### 4.2.2 Distribution, sample and provenance

**The model scores on the mean**, for comparability: a single figure per interval, so two assessments and two organisations can be compared without renegotiating the statistic. That is a deliberate simplification and it hides the tail, which is where response actually fails.

Where the assessor has them, the following are recorded and reported alongside the mean, and none of them changes the score:

| Field | Why it is reported |
|---|---|
| **P50** | Shows whether the mean is being dragged by outliers. |
| **P90** | The number that decides whether you contain in time. A mean that clears the breakout window with a P90 that does not is a failing capability with a passing average. |
| **Sample size** | How many timed incidents the figures were drawn from. |
| **Measurement window** | The from and to dates the sample covers, so a reader can tell a current figure from a historical one. |
| **Severity tier** | Which incidents were counted. Tempo measured across all severities is not comparable with tempo measured on crown-jewel incidents. |
| **Data source** | Where the timestamps came from — case system, SIEM, exercise records, manual reconstruction. |

**Below ten timed incidents, the mean is flagged as underpowered.** Ten is the point below which a mean stops describing a distribution and starts describing an anecdote: with a handful of incidents the figure is dominated by whichever one went worst or best, and a P90 cannot be estimated from four points at all.

The flag is a **caveat on the reading, not a ceiling on the band**. It is reported, not enforced. The estimated case is already governed by R5b, and stacking a second penalty on small samples would punish the organisations that have started measuring rather than the ones that have not.

#### 4.2.3 Breakout source hierarchy

Breakout time is the denominator of every tempo figure the model reports, so where it came from decides how much the result can carry. Five ranks, of which the first two count as **measured in the assessed environment** and the remaining three are **estimates**:

| Rank | Source | Measured? | What it means |
|---|---|---|---|
| **1** | Observed in our own incidents or exercises | **Measured** | Timed from your own incident or exercise records. The strongest basis, and the only one that supports a validated claim without further caveat. |
| **2** | Actor-specific threat intelligence | **Measured** | A named-actor figure from intelligence covering the adversaries you have prioritised. Cite the report and its date. |
| **3** | Sector-specific evidence | Estimate | A figure for organisations like yours, but not for your estate. Directionally useful; not a measurement of you. |
| **4** | Published industry benchmark | Estimate | A cross-industry average such as a vendor threat report. A sensible default when nothing better exists, and explicitly an estimate. |
| **5** | Assessor estimate | Estimate | A professional judgement with no published basis. Record the reasoning; it cannot support a claim above L3. |

An unrecorded provenance is treated as rank 5. This is deliberate: silence about where a number came from is not evidence that it was measured.

**Recording requirement.** Every breakout figure must carry three fields: its **source rank**, its **reference** (the report, the incident set or the reasoning), and its **date**. A figure without them is unauditable, and a tempo result computed from an unauditable denominator is a decoration.

Note that *published* and *estimated* are different claims and only the second one governs R5b. A published cross-industry average has a citable source rather than being invented — and it is still an estimate for a specific estate. Both statements are true at once, and conflating them is how a borrowed number becomes a validated capability.

#### 4.2.4 Interpretation

| Tempo Ratio | Meaning |
|---|---|
| TR < 0.5 | Containment lands well inside breakout — genuine tempo advantage |
| 0.5 ≤ TR < 1.0 | Contains before spread, with little margin |
| 1.0 ≤ TR < 2.0 | Structurally behind — you contain a spread intrusion, not a foothold |
| TR ≥ 2.0 | You are performing post-incident recovery and calling it response |

Tempo Ratio drives constraint **R5**; provenance drives **R5b** (§8.5).

**Malformed tempo is not the same as absent tempo.** A partially completed tempo entry, a non-positive breakout time, or detect, decide and contain all recorded as zero, are rejected as measurements and reported as malformed. Three zeros are an empty form that happens to parse, not instantaneous containment; left unguarded they produce a Tempo Ratio of 0.00 and lift a perfect assessment to L5. A breakout time alone, with no intervals, is not a failed tempo measurement — it comes from picking a priority actor — and is reported as tempo not supplied.

---

## 5. Maturity bands

Six bands, identical arithmetic ranges to TID-CMM so the two scores sit on one axis.

| Band | Range | Name | Definition |
|---|---|---|---|
| **L0** | 0.00–0.99 | **Improvised** | Response is individual heroics. The outcome depends on who happens to be on shift. This is a build project, not an improvement project. |
| **L1** | 1.00–1.99 | **Documented** | Plans exist on paper. Execution is unrehearsed and personality-dependent. The document has never met a clock. |
| **L2** | 2.00–2.99 | **Repeatable** | Playbooks and tooling are in place and consistently executed for known scenarios, but remain unproven under time pressure. **The most common band, and the band where response spending most exceeds response capability.** |
| **L3** | 3.00–3.99 | **Threat-informed** | Response traces to modelled attack paths and crown jewels. Containment is pre-authorised. Exercises are regular and evidenced. **The realistic target for most organisations.** |
| **L4** | 4.00–4.99 | **Validated** | A standing validation function exists. Containment is demonstrably inside the breakout window for priority actors. Requires tempo evidence — an L4 claim without measured MTTDecide is not assessable. |
| **L5** | 5.00 | **Adaptive** | Response adapts as the threat model changes. Treat any L5 claim with scepticism unless the evidence is exceptional. |

**Band gate:** L4 and L5 are unreachable without supplied tempo evidence, and unreachable on estimated tempo (§8.5, §8.5.1). L3 is unreachable with any T1-tier blind cell (§8.6).

---

## 6. The eight domains and 58 sub-capabilities

Weights total 100. Sub-capability counts total 58, matching TID-CMM's structure so the two radars superimpose.

| ID | Domain | Weight | Subs | Core question |
|---|---|---|---|---|
| **RP** | Response Preparation & Readiness | 10% | 7 | Are we set up to run an incident at all? |
| **RA** | Response Authority & Decision Rights | 12% | 6 | Is someone permitted to act, in time? |
| **RE** | Response Engineering & Playbooks | 16% | 10 | Is response engineered, or written? |
| **CE** | Containment, Eradication & Recovery | 14% | 8 | Do we have graded options that actually work? |
| **AO** | Automation & Orchestration | 12% | 7 | Does machine speed actually reach the decision? |
| **FI** | Forensics, Evidence & Investigation | 10% | 6 | Do we know what actually happened? |
| **RV** | Response Validation & Exercising | 14% | 7 | Have we proven any of this? |
| **RG** | Response Governance, Metrics & Improvement | 12% | 7 | Does the capability compound? |

Every sub-capability is scored 0–5. Sub-weights are within-domain and total 100 per domain.

The sub-capability name, weight, level-3 anchor and evidence descriptor in each table below are reproduced from the machine-readable model and are canonical there. The score-0 and score-5 columns are descriptive anchors published in this document only.

---

### RP — Response Preparation & Readiness · 10%

*Whether the organisation can start an incident cleanly. Cheap to build, catastrophic to lack, and consistently over-scored.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **RP-1** | Incident response plan and scope | 15 | No plan, or a template never adapted | Plan is organisation-specific, names roles, covers cloud/identity/OT as applicable, reviewed annually | Plan is versioned, scenario-indexed to the lattice, and demonstrably used in the last three incidents | Plan document + version history + incident references |
| **RP-2** | Severity and classification model | 15 | Ad-hoc severity by opinion | Documented severity matrix driven by crown-jewel and business impact, applied consistently | Severity assignment is auditable, drives automatic authority and comms activation, and is reviewed against outcomes | Severity matrix + sample of classified incidents |
| **RP-3** | Roles, rotas and 24×7 reachability | 15 | Best-effort; unclear who responds out of hours | Named IR roles, documented rota, tested reachability, defined deputies | Reachability is tested unannounced; coverage gaps are measured and remediated | Rota + unannounced call-out test records |
| **RP-4** | Response tooling readiness | 15 | Tools are assumed available, never checked | Response tooling (EDR console, SOAR, forensic kit, out-of-band comms) is inventoried and access-tested | Tooling readiness is continuously monitored; break-glass access is tested and time-bounded | Tool inventory + access test log |
| **RP-5** | Out-of-band communications and war-room | 10 | None; incident comms would run on the compromised estate | Documented out-of-band channel and bridge, contact tree maintained | Out-of-band channel is exercised at least annually under assumed-compromise conditions | Exercise report showing out-of-band use |
| **RP-6** | Third-party, retainer and supplier readiness | 15 | No retainer; no supplier response terms | IR retainer or in-house equivalent in place, SLAs known, onboarding pre-completed | Retainer is exercised; critical suppliers have tested notification and joint-response paths | Retainer contract + activation or exercise record |
| **RP-7** | Legal, regulatory and communications preparation | 15 | Unclear obligations; no prepared position | Reporting obligations mapped (GDPR/DORA/NIS2/sector), counsel identified, holding statements drafted | Regulatory clocks are automated into the incident process and have been met under exercise or real conditions | Obligation register + evidence of clock adherence |

---

### RA — Response Authority & Decision Rights · 12%

*The domain no existing maturity model measures. Response fails here more often than it fails on tooling.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **RA-1** | Pre-authorised containment actions | 25 | Every containment action requires case-by-case approval | A defined set of containment actions is pre-authorised by asset class and severity, documented and signed off by business owners | Pre-authorisation covers every T1/T2 lattice cell, is reviewed as the estate changes, and is exercised | Signed authorisation matrix mapped to lattice cells |
| **RA-2** | Decision latency measurement (MTTDecide) | 20 | Not measured; folded into MTTR | MTTDecide is measured separately from MTTR, per severity, and reported | MTTDecide is measured, targeted against breakout time, and trending against an explicit objective | Metric series with timestamps from the case system |
| **RA-3** | Escalation thresholds and triggers | 15 | Escalation is by instinct | Documented, objective escalation triggers tied to severity and crown-jewel involvement | Triggers fire automatically from the case system; false-escalation and missed-escalation rates are tracked | Trigger definitions + escalation audit trail |
| **RA-4** | Out-of-hours and degraded-mode authority | 15 | No authority exists out of hours | Named out-of-hours decision authority with documented deputies and a defined maximum response time | Out-of-hours authority is tested unannounced and meets its time objective | Unannounced out-of-hours exercise record |
| **RA-5** | Business impact acceptance and risk ownership | 15 | Security cannot accept the business impact of containment; actions stall | Business owners have accepted, in advance and in writing, the disruption cost of defined containment tiers | Acceptance is reviewed against realised impact after incidents and exercises; disputes have a defined arbiter | Signed impact-acceptance records |
| **RA-6** | Crisis and executive decision structure | 10 | No crisis structure; ad-hoc escalation to whoever answers | Documented crisis management structure with defined activation criteria and executive decision rights | Crisis structure has been activated and evaluated; executives have participated in exercises within twelve months | Crisis activation log + executive exercise record |

---

### RE — Response Engineering & Playbooks · 16%

*The largest domain, mirroring TID-CMM’s Detection Engineering weight. Response content deserves the same engineering discipline as detection content.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **RE-1** | Playbook coverage of the lattice | 15 | Playbooks exist for a handful of familiar scenarios | Every Tier-1 and Tier-2 lattice cell has an associated playbook or documented response procedure | Full lattice coverage including T3; gaps are tracked as a managed backlog with owners and dates | Playbook-to-cell coverage map |
| **RE-2** | Threat-informed derivation | 12 | Playbooks are generic or vendor-supplied | Playbooks derive from modelled attack paths, crown jewels and priority actor TTPs | Derivation is traceable end to end: actor → path → stage → asset → playbook → action | Traceability matrix: actor → path → stage → asset → playbook |
| **RE-3** | Playbook-as-code and version control | 12 | Playbooks live in documents or a wiki, unversioned | Playbooks are stored in version control with review, history and change approval | Playbooks are machine-readable, linted in CI, and deploy to the orchestration platform from the repository | Repository + CI pipeline evidence |
| **RE-4** | Playbook testing before release | 12 | Never tested before use | Every playbook is executed in a test or staging context before production release | Automated playbook regression testing runs on change; failures block release | CI test results |
| **RE-5** | Parameterisation and reusability | 8 | Each playbook is bespoke prose | Playbooks are built from reusable, parameterised response actions rather than duplicated steps | A response-action library exists (RE&CT/D3FEND-aligned) and playbooks compose from it | Response-action library + composition evidence |
| **RE-6** | Decision points and human gates | 10 | Playbooks assume a single path with no branch logic | Playbooks contain explicit decision points, named decision owners, and defined default actions on timeout | Decision points carry measured latency data and are optimised against tempo objectives | Playbook with instrumented decision points |
| **RE-7** | Playbook lifecycle and deprecation | 8 | Playbooks accumulate; nothing is retired | Playbooks have owners, review cycles and a deprecation process | Playbook health is measured (staleness, execution success rate, drift against estate) | Lifecycle register + health metrics |
| **RE-8** | Response action mapping to ontology | 8 | No mapping | Response actions are mapped to RE&CT RA-codes and/or D3FEND techniques | Mapping is bidirectional and used to identify uncovered response actions against modelled techniques | Mapping export |
| **RE-9** | Cross-domain response content | 8 | Endpoint-only playbooks | Playbooks exist for identity, cloud control plane and SaaS compromise, and OT where applicable | Cross-domain playbooks are rehearsed and reflect the specialist containment primitives of each estate | Domain-specific playbooks + exercise evidence |
| **RE-10** | Response content sharing and reuse | 7 | Nothing shared internally or externally | Playbooks and actions are shared across teams and regions against a common standard | Organisation contributes to or consumes from open response content communities with a defined ingestion process | Shared repository or contribution record |

---

### CE — Containment, Eradication & Recovery · 14%

*Graded, reversible options with known blast radius. Maps to D3FEND Isolate, Evict and Restore.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **CE-1** | Tiered containment options per asset class | 16 | One blunt option, or none | Each asset class has graded containment tiers (observe / restrict / isolate / disable) with documented business impact | Tiers are exercised, impact is measured against prediction, and selection is guided by the severity model | Containment tier matrix + exercise data |
| **CE-2** | Blast radius control and reversibility | 14 | Containment effects are unknown and irreversible | Blast radius is documented per action; every action has a defined rollback | Rollback is tested; unintended-impact rate is measured and trending down | Rollback test records |
| **CE-3** | Identity containment (Tier-0) | 14 | No ability to contain identity compromise at speed | Session revocation, credential reset, token invalidation and trust-break procedures exist and are owned | Identity containment is proven inside the breakout window and covers federated and non-human identities | Exercise record with timings |
| **CE-4** | Network and egress containment | 10 | No dynamic blocking capability | Egress blocking, sinkholing and segment isolation are available and documented | Network containment is automated where safe, with proven propagation times | Automation + propagation timing evidence |
| **CE-5** | Eradication completeness and re-entry prevention | 14 | Eradication means "we rebuilt the box" | Eradication addresses persistence, credentials and access paths, with a documented completeness checklist | Re-infection rate is measured; eradication is verified by independent scoping before closure | Eradication verification records + reinfection metric |
| **CE-6** | Recovery, restoration and integrity verification | 12 | Recovery is IT's problem, uncoordinated with IR | Recovery procedures are integrated with IR, with defined RTO/RPO for crown jewels and integrity verification before return | Recovery is exercised at crown-jewel scale; restored systems are verified clean against a defined standard | Recovery exercise report |
| **CE-7** | Backup resilience against destructive attack | 10 | Backups are online and reachable with production credentials | Immutable or logically isolated backups exist for crown jewels, with separate credentials | Restore from immutable backup is tested at scale within the recency window, with measured time-to-restore | Restore test record at scale |
| **CE-8** | Degraded-mode and business continuity operation | 10 | No defined degraded mode | Documented degraded-mode operation for crown-jewel services, agreed with the business | Degraded mode is exercised with business participation and meets agreed service floors | Continuity exercise record |

---

### AO — Automation & Orchestration · 12%

*Machine speed is only useful if it is permitted to reach the decision. Constrained by authority — see R2.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **AO-1** | Enrichment and triage automation | 14 | All enrichment is manual | Automated enrichment of alerts (asset, identity, intel, prior cases) before analyst contact | Enrichment quality is measured; analyst time-to-context is tracked and improving | Time-to-context metric series |
| **AO-2** | Automated containment coverage | 20 | No automated containment | Automated containment exists for defined low-risk, high-confidence scenarios with pre-authorisation | Automated containment covers the majority of T1/T2 cells where safe, with measured execution success | Automation coverage map + execution success rate |
| **AO-3** | Human-in-the-loop gate design | 14 | Either everything is manual or automation runs unsupervised | Gates are explicitly designed: which actions require a human, which do not, and what happens on timeout | Gate placement is reviewed against measured decision latency and incident outcomes | Gate design document + latency data |
| **AO-4** | Integration coverage and API depth | 14 | Few integrations; response requires console-hopping | The orchestration platform integrates with the tools controlling every in-scope asset class | Integration health is monitored; failures alert; coverage tracks estate change automatically | Integration inventory + health monitoring |
| **AO-5** | Automation reliability, safety and rollback | 14 | Automation fails silently or has caused unplanned outages | Automation has error handling, safety limits (rate and scope caps) and rollback paths | Automation failure rate and unintended-impact rate are measured; safety limits are tested | Reliability metrics + safety test evidence |
| **AO-6** | Case management and workflow integrity | 14 | Incidents tracked in email or chat | A case system holds all incidents with structured timeline, actions and timestamps | Case data is complete enough to compute MTTD/MTTDecide/MTTC automatically without manual reconstruction | Automated derivation of MTTD/MTTDecide/MTTC from case data |
| **AO-7** | Automation change control | 10 | Automation is edited live in production | Automation changes go through review and testing before deployment | Automation is deployed via pipeline with staged rollout and automated regression testing | CI/CD evidence |

---

### FI — Forensics, Evidence & Investigation · 10%

*Scoping accuracy determines eradication completeness. Under-invested almost universally.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **FI-1** | Triage and scoping capability | 22 | Scoping is guesswork; "patient zero" is rarely established | Structured triage produces a defensible scope: affected assets, identities and timeframe | Scoping accuracy is measured retrospectively; under-scoping incidents are treated as findings | Scope-accuracy review records |
| **FI-2** | Evidence acquisition capability | 18 | No acquisition capability; systems are wiped before analysis | Volatile and disk acquisition is possible for each in-scope asset class within a defined time | Acquisition is proven at speed for cloud and ephemeral workloads, not only endpoints | Acquisition test records including cloud and ephemeral workloads |
| **FI-3** | Chain of custody and evidence handling | 15 | No chain of custody | Documented chain of custody suitable for internal disciplinary and insurance purposes | Evidence handling meets a standard suitable for legal or regulatory proceedings and has been reviewed by counsel | Custody procedure + legal review |
| **FI-4** | Timeline reconstruction | 15 | No timeline produced | Incident timelines are produced for significant incidents with source attribution | Timeline construction is partly automated from case and telemetry data and is used to compute tempo metrics | Sample timelines + automation |
| **FI-5** | Log and evidence retention adequacy | 15 | Retention is shorter than typical dwell time | Retention for crown-jewel-relevant sources exceeds median dwell time for priority actors | Retention is set from measured dwell time and validated by successful historical investigation | Retention policy + successful historical lookback case |
| **FI-6** | Malware and artefact analysis access | 15 | No capability, internal or external | Analysis capability available in-house or via retainer, with defined turnaround | Analysis output feeds detection engineering and playbook updates within a defined cycle | Analysis reports + resulting detection or playbook changes |

---

### RV — Response Validation & Exercising · 14%

*The ceiling-setting domain. Every other domain is capped at RV + 1 by constraint R1.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **RV-1** | Tabletop exercise programme | 12 | None, or one-off compliance theatre | Regular tabletops covering crown-jewel scenarios with the right participants, findings tracked | Tabletops are scenario-derived from the lattice, include executives and third parties, and drive tracked remediation | Exercise reports + remediation tracking |
| **RV-2** | Technical live-fire exercising | 22 | Never; playbooks have never been executed under exercise conditions | Playbooks are executed technically against emulated adversary activity in a controlled window | A standing programme executes live-fire against production or production-equivalent, on a defined rotation across the lattice | Live-fire schedule + results per lattice cell |
| **RV-3** | Purple team and end-to-end response validation | 18 | Purple team ends at detection | Purple team exercises continue through containment and eradication, not stopping at the alert | Every priority attack path is validated end to end — detect, decide, contain, evict, restore — on a defined cycle | End-to-end purple team reports |
| **RV-4** | Containment timing measurement under exercise | 16 | No timings captured | Exercises capture MTTD, MTTDecide and MTTC and compare them to breakout time | Timing objectives are set per stage, measured every exercise, and drive engineering priorities | Timing dataset across exercises |
| **RV-5** | Validation recency and coverage tracking | 12 | Not tracked | Validation status and date tracked per lattice cell, with expiry applied | Expiry is automated; a validation-debt figure is reported to leadership alongside coverage | Automated validation register + validation-debt figure |
| **RV-6** | Failure injection and response-function resilience | 10 | Never tested | Response is exercised under degraded conditions (key tool unavailable, key person absent, SIEM down) | Degraded-mode response is routinely injected and the function meets defined floors | Degraded-mode exercise records |
| **RV-7** | Independent assessment and red team | 10 | No independent challenge | Periodic independent red team or assessment includes response effectiveness, not just breach success | Independent assessment is scoped explicitly against TIR-CMM lattice cells and its findings are tracked to closure | Independent report + closure evidence |

---

### RG — Response Governance, Metrics & Improvement · 12%

*Whether the capability compounds. Closes UTIOM’s Kaizen loop and feeds findings back into TID-CMM.*

| ID | Sub-capability | Wt | Score 0 | Score 3 | Score 5 | Evidence for 4–5 |
|---|---|---|---|---|---|---|
| **RG-1** | Response metrics programme | 18 | No metrics, or ticket counts only | MTTD, MTTDecide, MTTC, MTTR and containment success are measured and reported | Metrics are leading as well as lagging, tied to breakout time, and drive investment decisions | Metric definitions + reporting history |
| **RG-2** | Post-incident review discipline | 18 | Reviews happen rarely and blamefully | Blameless post-incident reviews occur for all significant incidents, with tracked actions | Reviews produce structural findings; action closure rate and recurrence rate are measured | PIR records + action closure metrics |
| **RG-3** | Feedback loop into detection (TID-CMM) | 16 | No feedback; response findings die in the report | Incident and exercise findings generate detection engineering work items | The loop is measured: proportion of incidents producing detection improvements, and time to deploy | Linked incident → detection change records |
| **RG-4** | Feedback loop into architecture and hardening | 12 | None | Findings generate hardening and architecture work items with owners | Architecture findings are tracked to completion and their effect is re-validated | Linked findings → architecture changes |
| **RG-5** | Response capability ownership and funding | 12 | No clear owner or budget for response capability | Named owner and dedicated budget for response engineering, distinct from SOC staffing | Investment is prioritised using measured gaps (TIR-CMM roadmap) rather than vendor cycles | Budget + prioritisation evidence |
| **RG-6** | Regulatory and executive reporting | 12 | Reporting is improvised under pressure | Defined reporting packs and regulatory notification procedures, with owners and clocks | Reporting has been executed under real or exercise conditions within the required clocks | Reporting evidence with timestamps |
| **RG-7** | Continuous improvement cadence | 12 | Improvement is project-driven and sporadic | A regular cadence reviews response capability against the model and adjusts the backlog | Improvement is continuous, measured against the lattice, and reassessment shows movement | Reassessment history |
---

## 7. Scoring mathematics

Identical in form to TID-CMM, so results are comparable and the workbook formulas transfer.

### 7.1 Sub-capability → domain

```
domain_score = Σ(sub_weight × sub_score) / Σ(sub_weight)
```

Not-applicable sub-capabilities are removed from both numerator and denominator. No penalty, no benefit.

An assessor declaring something out of scope is a judgement; never getting to it is not. Both are excluded from the mean, so the report distinguishes them: **not applicable** and **unanswered** are counted and reported separately, alongside the proportion of total domain weight actually covered.

### 7.2 Domain → overall

```
overall_score = Σ(domain_weight × domain_score) / Σ(domain_weight)
```

Two behaviours of this formula are load-bearing and were previously documented only in the engine.

**A domain with no answers is excluded and the remaining weights are renormalised. It is not scored zero.** The denominator is the sum of the weights of the domains that carry an answer, not 100. An unmeasured domain neither helps nor hurts the overall figure.

This is not a convenience. Scoring an unanswered domain as 0.00 would be inconsistent with how an unanswered sub-capability is handled one level down — excluded from numerator *and* denominator — and it would be actively dangerous. A zero in RV becomes a ceiling of 1.00 under R1, and a zero in RA becomes a ceiling of 1.00 under R2; either would drag every other domain to 1.00, so a half-finished assessment would read as a catastrophic organisation rather than an incomplete one. The report states which domains were unscored and what proportion of the model's weight the result actually covers, so a partial assessment is legible as partial.

**Pulse depth bypasses the R3 evidence cap.** Pulse does not collect evidence levels at all, so applying an evidence cap to it would penalise every answer as though evidence had been sought and found absent. It was never sought. **Pulse's L3 band ceiling under R7 carries that weight instead**, which is the honest place for it: the reason a Pulse result cannot substantiate a high claim is the depth of the assessment, not the grade of evidence behind any individual answer.

The bypass applies to Pulse and to nothing else. At Baseline and Assurance depth a blank evidence level is treated as VC0 and caps the score at 1.

### 7.3 Lattice → reported coverage

The lattice does **not** feed the overall domain score arithmetically. It is reported alongside it, exactly as TID-CMM reports Validated Coverage Score alongside domain maturity. Conflating them would let broad shallow coverage mask a structural inability to act.

The lattice does, however, bind the score through constraint **R6** (§8.6) and through prioritisation (§9).

### 7.4 Reported output set

| Output | Form |
|---|---|
| Overall maturity | 0.00–5.00 + band L0–L5 |
| Domain scores | 8 × 0.00–5.00, radar chart |
| Coverage | Subs scored / N/A / unanswered; domains scored; proportion of weight covered |
| VRS_cj | 0–100% |
| Engineered rate / Proven rate / Declared proven rate | 0–100% each — the headline gap, and the validation debt |
| Blind cells (T1/T2) | Count + named list |
| Recency expiries | Per cell and per sub-capability, with the reason each expired |
| Containment Margin per priority actor | Minutes, signed |
| Tempo Ratio per priority actor | Ratio, with provenance and sample fields |
| Constraint adjustments | Self-assessed vs adjusted, per constraint, independent and marginal effect |
| Band ceilings | Every ceiling considered, each marked binding or not binding |
| Consistency checks | Informative findings (§14) — never a score change |
| Ranked roadmap | Ordered list with impact scores |

---

## 8. Integrity constraints

TID-CMM's central innovation is that constraints apply **mechanically**, lower scores only, and cannot be argued with. TIR-CMM inherits this exactly and adds the constraints unique to response.

**Application order:**

```
R3  (sub-capability level, evidence — including §3.6 evidence recency)
 ↓
compute domain scores  (unanswered domains excluded, weights renormalised)
 ↓
R4  (detection dependency, cross-model)
 ↓
R2  (authority ceiling)
 ↓
R1  (rehearsal ceiling)
 ↓
compute overall score
 ↓
R5 / R5b (tempo band caps)   R6 (blind-cell band cap)   R7 (depth and governance)
```

Each constraint operates on figures already adjusted by the preceding ones. §3.6 recency is applied to the lattice before any lattice metric is computed, and to evidence levels before R3 caps.

**Reporting rule — independent effect, not marginal effect.** Because R1 is applied last and is usually the tightest ceiling, it routinely subsumes R2 and R4, which would then appear in the report as "no effect" even where they are diagnostically the most important finding. The tool must therefore compute and display each constraint's **independent** effect — what it would cap, evaluated against the pre-constraint domain scores — alongside its marginal effect on the final number. Without this, R2 and R4 become invisible in exactly the organisations they were written for.

**Band-ceiling rule — evaluate against the earned band.** Every band ceiling is evaluated against the band the *score* earns, not against the running band after other ceilings have applied. Testing against the running band caused a ceiling landing on a band that a previous ceiling had already reached to be discarded as "not binding", so a Tier-1 blind cell (R6, L2) vanished from the report whenever tempo (R5) had already capped to L2 — exactly the co-occurring failure the model exists to surface. Every ceiling considered is recorded, each marked binding or not binding.

**A ceiling derived from an unmeasured domain is not a ceiling of zero.** Where RV or RA is unscored, R1 or R2 respectively cannot be evaluated and is reported as unevaluated, not as a cap of 1.00.

### 8.1 R1 — Rehearsal Ceiling

```
∀ domain d ≠ RV:   d ≤ RV + 1
```

> **An unrehearsed playbook is an assumed capability.**

The direct analogue of TID-CMM's C1. Response is more exposed to this than detection: a detection rule at least executes automatically every day, whereas a playbook that is never run may not survive first contact with the tooling, the people or the clock.

### 8.2 R2 — Authority Ceiling

```
AO ≤ RA + 1
CE ≤ RA + 1
```

> **Automation you are not permitted to fire is a demonstration, not a capability.**

The constraint with no equivalent in any existing model. It expresses the single most common structural failure in incident response: a technically capable organisation that cannot act because nobody with authority is available, informed, or willing to accept the business impact. Buying a better SOAR does not move this score.

### 8.3 R3 — Evidence Rule

```
sub_score := min(sub_score, cap(evidence_level))
cap = { VC0: 1, VC1: 2, VC2: 4, VC3: 5 }
a blank evidence level is treated as VC0
VC3 with a missing or >12-month-old date grades to VC2 first (§3.6.2)
```

> Unsupported claims must not outrank evidenced assessments.

**The cap is a ceiling, not a conversion.** A high evidence grade *permits* a high score; it does not create one. VC3 evidence behind a capability you score 2 leaves the score at 2.

| Level | Class | Cap | What the assessor must be able to produce |
|---|---|---|---|
| VC0 | Assertion only | 1 | None, or a verbal claim |
| VC1 | Design or policy | 2 | A document, policy, screenshot, or partial implementation |
| VC2 | Implemented and tested | 4 | Implementation evidence plus a test result, ticket, or system record |
| VC3 | Repeatably validated | 5 | Recent, repeatable, independently reviewable proof from the live environment |

The artifact must be named and dated: an incident ticket ID, an exercise report with a date, a signed authorisation matrix, a timestamped containment log. "We do that" is not an artifact, and from v1.0 an undated artifact is not a VC3 one.

R3 is applied at Baseline and Assurance depth. Pulse bypasses it for the reason given in §7.2.

### 8.4 R4 — Detection Dependency (cross-model)

```
RE ≤ D + 1
CE ≤ D + 1
FI ≤ D + 1
```

> **You cannot respond to what you never saw.**

This is the structural hinge that makes TIR-CMM complete TID-CMM rather than duplicate it. It also creates the correct incentive: an organisation cannot reach a high response score by buying response tooling while remaining blind.

#### 8.4.1 What `D` is — settled

**`D` is TID-CMM's constraint-adjusted, pre-substitution overall detection score.** Three words in that sentence are load-bearing.

**Constraint-adjusted.** `D` is the figure after TID-CMM's own constraints have run, not its self-assessed number. TIR-CMM consumes a detection score that has already been made honest by its own model; re-deriving it here would be duplication, and consuming the self-assessed figure would import the optimism TID-CMM exists to remove.

**Pre-substitution.** Under the bridge (§11), TID-CMM substitutes its IR domain with a TIR-CMM overall score where one exists. `D` must be taken **before** that substitution. Otherwise TIR-CMM's own result re-enters TIR-CMM through the detection-to-response interface: a high response score raises TID-CMM's IR domain, which raises TID-CMM's overall, which raises `D`, which lifts the R4 ceiling on RE, CE and FI, which raises the response score. That is circular inflation, and the ordering is what prevents it. The import contract carries the pre-substitution figure as an explicit field for this reason.

**Overall.** `D` is the overall detection score, not the telemetry or data-collection domain.

**The telemetry domain is not used as a second cap.** The argument for using it is real — response depends most acutely on visibility, and a detection programme strong on content and thin on collection fails on exactly the asset classes response needs. The argument against it is decisive: TID-CMM has already applied its own visibility constraints to that weakness, and it is already reflected in the overall score `D` is drawn from. Capping again on the same weakness penalises it twice, once inside TID-CMM's arithmetic and once inside TIR-CMM's, and double-counting a single deficiency is exactly the behaviour the independent-effect reporting rule exists to expose elsewhere in this model.

The telemetry figure is instead **displayed as diagnostic context**. Where it sits materially below the overall detection score — a full point or more — the assessment raises a review warning (consistency check **C13**, §14) rather than a second penalty. The warning names the gap, states that R4 has already applied the single detection ceiling, and asks the assessor to check whether the telemetry gap falls on the same asset classes as the Tier-1 lattice cells. That is a diagnosis. It is not arithmetic, and it never changes a score.

#### 8.4.2 Sources of `D`

**TIR-CMM does not require TID-CMM.** `D` may come from any of three sources, each justifying a different band ceiling:

| Source of `D` | Band ceiling | Rationale |
|---|---|---|
| **Imported TID-CMM assessment** | none — assessable to L5 | Detection maturity is independently evidenced. |
| **Built-in prerequisite check** (§8.4.3) | **L4** | A self-assessment is sufficient to assess response up to validated maturity. Claiming *adaptive* response maturity on an unaudited detection claim is not credible. |
| **Nothing supplied** | **L3** | The response score rests on an unaudited claim, and the report says so. |

The point of the middle row is that an organisation with no detection maturity model of any kind is a normal customer for this model, not a rejected one. Turning it away would be both unhelpful and self-defeating: response is precisely where such an organisation is most exposed.

Where `D` is absent entirely, R4 is reported as **unevaluated** — the constraint cannot be computed, which is a different statement from "it did not bind".

#### 8.4.3 The prerequisite check

Eight questions, scored 0–5, of which at least six must be answered for the proxy to count.

**Detection (produces `D`):**

| ID | Question |
|---|---|
| PD-1 | Telemetry coverage of crown jewels |
| PD-2 | Where detection content comes from |
| PD-3 | Detection testing |
| PD-4 | Alert quality reaching a responder |
| PD-5 | Identity and cloud detection coverage |

**Threat modelling (produces `TM`, governing confidence in lattice scoping):**

| ID | Question |
|---|---|
| PM-1 | Crown jewel definition |
| PM-2 | Attack path modelling |
| PM-3 | Adversary prioritisation |

The modelling questions are *scored*, not assumed. A declared attack path from an organisation that has never modelled one does not earn the same scope quality as an imported modelled path, and PM-2 is where that difference is recorded.

---

## 8A. Readiness lenses

The eight domains describe how the capability is **built**. Leadership asks a different question, and it is the one that decides whether the assessment gets acted on:

| Lens | Question | Subs |
|---|---|---|
| **Readiness to respond** | If it started now, could we act on it in time? | 27 |
| **Readiness to recover** | Could we get the business back, clean, and prove it? | 14 |
| **Operational resilience** | Could we keep running through it, and be harder to hit next time? | 15 |

Each lens re-cuts the same 58 sub-capabilities, with a weight of 2 for sub-capabilities core to that outcome and 1 for contributing ones.

**They are lenses, not partitions.** Eight sub-capabilities serve more than one outcome, which is correct — CE-7 (backup resilience) is both a recovery capability and a resilience one, and forcing it into a single bucket would misrepresent it. Ten sub-capabilities sit in no lens, because they describe how the capability is engineered rather than what outcome it delivers.

Each lens reports two figures:

- **Claimed** — the weighted mean of its R3-adjusted sub-capability scores.
- **Proven** — the same figure with the rehearsal ceiling (RV + 1) applied, exactly as it is applied everywhere else in the model.

**Both are reported with equal prominence, and this is deliberate.** When the rehearsal ceiling binds — which is the common case — it flattens all three *proven* figures to the same number, destroying the lenses' only job of discriminating between the three outcomes. The *claimed* figure is what still separates them, and the gap between claimed and proven is itself the most useful output.

*Note:* R4 also subsumes the role of TID-CMM's C4 (intent ceiling) for the response module, since the threat modelling that produces the lattice is imported from TID-CMM's TM domain rather than re-assessed here.

### 8.5 R5 — Tempo Ceiling

Applies to the **band**, not to domain scores.

```
TR = (MTTD + MTTDecide + MTTC) / B(a_priority)

TR ≥ 2.0            ⟹  band capped at L1
1.0 ≤ TR < 2.0      ⟹  band capped at L2
tempo not supplied  ⟹  band capped at L3, report marked "tempo unverified"
```

> **A response capability structurally slower than the adversary is not a level 3 capability, whatever the documentation says.**

This is the constraint that prevents TIR-CMM from becoming a paperwork audit. An organisation can hold excellent playbooks, full automation coverage and a rehearsal programme, and still be losing every race. The tempo cap makes that visible in one number that a board understands.

#### 8.5.1 R5b — Estimated Tempo Ceiling

```
breakout window estimated  ⟹  band capped at L3
interval timings estimated ⟹  band capped at L3
```

> **A measured-looking ratio computed from a borrowed number is not a measurement.**

A tempo result computed from an **estimated breakout window**, or from **estimated interval timings**, caps the band at L3 regardless of how favourable the ratio is. L4 and above require timings measured in the assessed environment.

**"Estimated" means not measured in the assessed environment.** It does not mean invented, unsourced or careless. A published cross-industry benchmark from a reputable threat report is a legitimate default — it is the right thing to use when nothing better exists, and the model ships such figures precisely so that an assessor is not left with an empty box. It is nevertheless an estimate *for a specific estate*, because the adversary a specific organisation actually faces sets that organisation's denominator, and a cross-industry average is by construction not that.

The mapping to the source hierarchy (§4.2.3) is mechanical: ranks 1 and 2 are measured; ranks 3, 4 and 5 are estimates; an unrecorded provenance is treated as an estimate. Interval timings are declared as estimated or measured by the assessor. If either side of the ratio is estimated, R5b applies, and the report states which side and why.

R5b interacts with R5 rather than replacing it. R5 tests whether the ratio is good enough; R5b tests whether the ratio means anything. A ratio of 0.30 built on an assessor's archetype passes R5 and is capped at L3 by R5b. A ratio of 5.08 built on measured timings fails R5 at L1, and R5b adds nothing.

**This rule was published as model behaviour from v0.2 and was not implemented until v1.0.** In v0.2 the archetype presets shipped with an `estimated` flag that nothing consumed: selecting a 1440-minute insider archetype produced a Tempo Ratio of 0.03 and permitted L5 on a guess. The rule described the intended behaviour correctly for a year; the engine did not perform it. It does now.

### 8.6 R6 — Blind-Cell Gate

```
∃ cell c : tier(c) = T1 ∧ RRS(c) = 0   ⟹  band capped at L2
```

> If there is a stage on a crown jewel where you have no response option at all, you are not threat-informed.

The response analogue of TID-CMM's structural-gap analysis, and consistent with UTIOM's Law 3 (crown jewels drive resource allocation) and Law 5 (blind spots are architectural choices).

R6 is evaluated on **effective** RRS, after §3.6. A cell can only reach status 0 by being scored 0; recency never drives a cell below 2, so recency can never create a blind cell.

### 8.7 R7 — Assessment Depth and Governance Ceiling

```
tier = Pulse                        ⟹  band capped at L3
tier = Baseline                     ⟹  band capped at L4
tier = Assurance ∧ ¬governance_core ⟹  band capped at L4
   where governance_core = GV-3 ∧ GV-4 ∧ GV-5
```

> The depth an assessment reaches, and whether it was independently calibrated, are themselves evidence about how far the result can be trusted.

R7 is the constraint that keeps the model honest about *itself*. A Pulse assessment asks twenty questions and reviews no evidence; whatever it returns, it cannot substantiate a threat-informed claim, so it caps at L3 and says so on the report. Baseline is self-service and self-graded, so it caps at L4 — L5 requires an assessment somebody else can check. Assurance removes the depth ceiling, but only if the governance core is in place: a scoring rubric applied consistently (GV-3), an independent review of the evidence (GV-4), and a declared conflict-of-interest position (GV-5).

Like every other band ceiling, R7 is **recorded only as binding when it binds**. A Baseline assessment scoring L2 is not "capped at L4" in any meaningful sense, and saying so would send the reader looking for a constraint that is not there. It is still listed among the ceilings considered.

### 8.8 Worked constraint effect

TID-CMM's published worked example moves an organisation from a self-assessed 2.44 to an adjusted 2.34 — a deliberately modest drop that demonstrates the mechanism without theatre. TIR-CMM behaves comparably on the decimal, and the honest shock arrives in the **band**, not the number.

### 8.9 Numerical validation — "Meridian Group"

The full pipeline is executed numerically against a fictional mid-size European financial services organisation (~9,000 staff) with a competent but optimistic security team: good preparation, real detection engineering investment, a well-regarded SOAR deployment, tabletop exercises — and no live-fire, no measured decision latency, nothing signed on containment authority.

This is a **validation trace against a fixture**, not a customer, not a pilot and not evidence about any real organisation. Every figure below is produced by running the shipped engine against the `demoState()` fixture, which is the same worked example the assessment tool loads. It is regenerated whenever the model changes, so it reproduces rather than being asserted.

**Structural checks:** domain weights sum to 100 ✓ · 58 sub-capabilities ✓ · all eight domains' sub-weights sum to 100 ✓

Detection maturity is imported at 2.34; the assessment is run at Baseline depth.

**Constraint trace:**

| Step | Effect |
|---|---|
| Self-assessed overall | **2.69** (L2) |
| R3 evidence rule | 6 sub-capabilities capped by their evidence level: RP-3 4→2, RP-4 4→2, RP-5 3→2, RA-1 4→2, RE-5 3→2, AO-5 3→2 |
| R4 detection dependency | ceiling 3.34 (detection 2.34 + 1) — does not bind, on its own or at final ordering |
| R2 authority ceiling | ceiling 3.05 (RA 2.05 + 1) — does not bind, on its own or at final ordering |
| R1 rehearsal ceiling | RV = 1.52 → ceiling 2.52. **Binds six of eight domains.** RP 3.20→2.52, RE 2.66→2.52, CE 2.68→2.52, AO 2.80→2.52, FI 2.85→2.52, RG 2.72→2.52 |
| Constraint-adjusted overall | **2.32** (L2), Δ −0.36 |

**Lattice:** 32 in-scope cells (inside the 24–38 design band) · VRS **45.8%** · VRS_cj **44.6%** · engineered rate **40.6%** · **proven rate 0.0%** against a declared proven rate of 3.1% · one blind Tier-1 cell: **S3 × A1 — privilege escalation on identity infrastructure, no response option at all**.

The proven rate is 0.0% because §3.6 recency now applies. The fixture's single Proven cell, S2 × A2, carries no proof date, so it reverts to Engineered with the reason recorded: *"No date recorded for the last proof, so it cannot be shown to be current."* This is the migration behaviour of §3.6.3 visible in the model's own worked example, and it is left visible deliberately rather than back-dated to keep the number pretty.

VRS_cj sits *below* VRS (44.6% against 45.8%), which is the signal worth acting on: this organisation's containment coverage is marginally thinner on the cells that matter most than it is on average.

The engineered-versus-proven gap (40.6% vs 0.0%) is the response analogue of TID-CMM's 48.9% vs 12.7%, and is proportionally wider — which is the expected and intended result, since response capability is rehearsed far less often than detection logic is exercised.

**Tempo:** breakout 62 min · MTTD 180 + MTTDecide 95 + MTTC 40 = 315 min · **Containment Margin −253 min** · **Tempo Ratio 5.08**. The fixture declares no breakout source, so provenance is treated as estimated and R5b is evaluated.

**Readiness lenses:** respond 2.53 claimed / 2.52 proven · recover 3.26 / 2.52 · resilience 2.30 / 2.30. The recover lens is where the rehearsal ceiling costs most: 0.74 of claimed readiness that nothing has ever tested — the signature of a mature disaster-recovery programme that has never been exercised as an incident.

**Band ceilings considered:**

| Ceiling | Band | Binding? |
|---|---|---|
| R5 — Tempo Ratio 5.08 ≥ 2.0 | L1 | **binds** |
| R5b — tempo rests on an estimated breakout window | L3 | considered, does not bind (sits above the earned band) |
| R6 — one Tier-1 blind cell | L2 | **binds** |
| R7 — Baseline depth | L4 | considered, does not bind |

Final reported band: **L1 — Documented**, against a self-assessed 2.69 that would have read L2 and, in most maturity models, been reported as "approaching level 3."

Both R5 and R6 are recorded as binding, because each is evaluated against the band the score earned (L2) rather than against the running band. A ceiling that sits above the achieved band, as R5b and R7 do here, is not holding the band down, and reporting it as though it were would misdirect the remediation effort.

**Mechanical roadmap, top five:** RV-2 live-fire exercising (100.0) · S3×A1 blind identity cell (100.0) · S0×A1 prevention and hardening on identity (82.1) · S0×A7 prevention and hardening on backup and recovery (82.1) · RA-2 decision-latency measurement (77.9).

Not one of the top five is a product purchase, and three of five are things this organisation believed it already had. That is the model working.

**Defects found and corrected by executing this validation rather than asserting it:** the roadmap scale mismatch (§9.3); the constraint-reporting problem (§8, reporting rule); a domain with no answers scoring 0.00 and becoming an R1 ceiling of 1.00 that dragged every other domain down (§7.2); band ceilings being recorded as binding when they sat above the achieved band; tempo figures that were absent, negative or all-zero being accepted as measurements; and, at v1.0, an estimated breakout window permitting L5 (R5b) and Proven cells that never expired (§3.6). None would have shipped noticed without running the arithmetic.

---

## 9. Prioritisation and roadmap generation

Two ranked lists are produced and then merged. Both are mechanical — no advocacy, no facilitator bias.

### 9.1 Capability gaps (sub-capability level)

Identical to TID-CMM's formula:

```
impact = (domain_weight / 100) × (sub_weight / 100) × gap_to_target × 1000
```

where `gap_to_target = max(0, 3 − effective_sub_score)`. The target is 3 — threat-informed — not 5.

### 9.2 Lattice gaps (cell level)

```
cell_impact = ω(c) × λ(stage) × (3 − RRS(c)) × 100

  ω = criticality tier weight (T1=3, T2=2, T3=1)
  λ = stage leverage (S0 1.6, S1 1.5 … S7 0.8)
```

This surfaces the correct counter-intuitive result: a status-0 cell at S1 on a Tier-0 identity asset outranks a status-2 cell at S7 on a peripheral system, even though the latter feels more urgent during an incident.

### 9.3 Normalisation before merge — mandatory

The two lists operate on incompatible scales.

| List | Maximum possible raw impact | Where it comes from |
|---|---|---|
| **Capability gap** | **92.40** | RV-2: domain weight 14, sub-weight 22, at a full three-point gap — (14/100) × (22/100) × 3 × 1000 |
| **Lattice cell** | **1440** | A T1 cell at S0 at status 0 — 3 × 1.6 × 3 × 100 |

Merging the raw figures produces a roadmap consisting **entirely of lattice cells**, with every capability gap buried. Each list is therefore normalised to its own maximum before merging:

```
normalised_impact = 100 × impact / max(impact in that list)
```

This is scale-free and survives any future reweighting. The merged list is then sorted descending, which produces a properly interleaved result.

### 9.4 Effort — what an item takes

Every sub-capability carries four further attributes, so the ranking can be turned into something a person can start on Monday:

| Attribute | Purpose |
|---|---|
| **act** | The concrete first action, imperative, one sentence. Not "improve decision latency measurement" but "measure the gap between alert-validated and containment-authorised timestamps on your last 20 incidents." |
| **win** | What you have once it is done — checkable, and usually the exact artifact constraint R3 demands for a score of 4 or 5. |
| **eff** | Effort tier: 1 = days, 2 = weeks, 3 = quarters. |
| **owner** | The role that must actually do it. For authority items this is a business or executive role, not security — which is the finding. |

The ranked roadmap is bucketed by **effort tier**:

| Bucket | Effort tier | Typical elapsed effort | Means |
|---|---|---|---|
| **Quick fixes** | 1 | 0–30 days | One person, no budget, no procurement. Writing things down, getting decisions made, measuring what you already do. |
| **Ninety-day moves** | 2 | 1–3 months | Coordination across teams, a scheduled exercise, or a contained piece of engineering. |
| **Structural work** | 3 | 6–12 months | Budget, procurement, hiring, architecture change or a standing programme. |

Effort spread across the 58: **20 / 23 / 15**.

The distribution is the argument. Authority and measurement work — the two domains that most often bind the score — is overwhelmingly effort tier 1. Five of the six RA sub-capabilities are quick fixes. This means the model's own ranking routinely puts *free* items at the top of the plan, and lets a CISO show a board that the largest single constraint on response capability costs a decision, not a purchase.

Lattice gaps are bucketed too: a **blind crown-jewel cell is a quick fix**, because naming an owner and writing an interim manual procedure is a days-not-quarters job and it is what lifts the R6 band cap. Every other lattice gap is a ninety-day move.

### 9.5 Scheduling — when an item is planned

The merged list is then sequenced into a plan:

1. Anything that lifts a **binding constraint** is promoted, because it unlocks score elsewhere. In practice, RA and RV work is almost always promoted — which is the model's central argument made arithmetically.
2. Items are grouped into **90-day / 180-day / 365-day** scheduling windows, matching UTIOM's roadmap tool.
3. The output names **capabilities and authorities required, not products** — inherited directly from TID-CMM's design principle.

### 9.6 Effort and scheduling are different concepts

§9.4 and §9.5 use different horizons and they are not two versions of the same thing. Conflating them is a reading error the earlier drafts of this specification invited, so the relationship is stated explicitly.

| | §9.4 effort | §9.5 scheduling |
|---|---|---|
| **Question answered** | How much does this item take? | When is this item planned to happen? |
| **Property of** | The item | The plan |
| **Horizons** | 0–30 days · 1–3 months · 6–12 months | 90 days · 180 days · 365 days |
| **Set by** | The model — each sub-capability carries a fixed effort tier | The assessor, from rank, dependency and capacity |
| **Changes between assessments?** | No | Yes, every time |

**Effort constrains scheduling; it does not determine it.** An effort-tier-1 item cannot be scheduled into the 365-day window on the grounds that it is hard, because it is not — but it can legitimately land there because it is ranked low or depends on something else. An effort-tier-3 item cannot be scheduled into the 90-day window, because the effort horizon says it will not land there whatever the plan claims.

The common and correct pattern is a 90-day window packed with effort-tier-1 items promoted for lifting a binding constraint, a 180-day window of tier-2 work, and a 365-day window holding tier-3 structural items plus whatever tier-1 and tier-2 work capacity pushed out. A 90-day window containing tier-3 items is a plan that will slip, and the model says so rather than accommodating it.

---

## 10. Standards crosswalk

TIR-CMM operationalises rather than replaces. Four ontologies anchor it, each doing a distinct job:

| Anchor | Role in TIR-CMM |
|---|---|
| **MITRE D3FEND**, ontology 1.5.0 (7 tactics; Isolate, Evict and Restore are response-side) | Structural mirror of ATT&CK. Provides the countermeasure vocabulary for containment primitives per asset class (§3.2) and the conceptual symmetry with TID-CMM. |
| **RE&CT** (6 response stages, RA1000–RA6000 series) | The operational spine. Provides the response-action library that RE-5 and RE-8 score against, and the phase structure practitioners already recognise. |
| **ATT&CK Mitigations (M-codes)** | Traceability bridge. Techniques scoped by TID-CMM carry M-codes; these map to lattice cells, so the response scope derives from the detection scope automatically. |
| **NIST SP 800-61r3** (April 2025) — *Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile* | Governance crosswalk. r3 abandons the four-phase lifecycle of r2 and organises incident response around CSF 2.0 Functions — GV, ID and PR as preparation, DE, RS and RC as incident response, with improvement fed back through ID.IM. TIR-CMM follows r3, not r2. |

### 10.1 Domain-level crosswalk

Identifiers are CSF 2.0 Categories, ISO/IEC 27035-1:2023 process phases, RE&CT stage code series, D3FEND tactics and SOC-CMM domains.

| TIR-CMM | NIST CSF 2.0 | SP 800-61r3 (CSF Functions) | ISO/IEC 27035-1:2023 | RE&CT | D3FEND | SOC-CMM | DORA / NIS2 |
|---|---|---|---|---|---|---|---|
| **RP** | GV.OC, PR.PS, RS.MA | Preparation (GV, ID, PR) | Plan and prepare | RA1000 | Model, Harden | Process, People | ICT continuity, incident mgmt |
| **RA** | GV.RR, RS.MA | Preparation (GV) | Plan and prepare | RA1000 | — | Business (Governance) | Management body accountability |
| **RE** | RS.AN, RS.MI | Respond (RS) | Detect and report; Respond | RA2000–RA5000 | Isolate, Evict | Process, Technology | Response & recovery plans |
| **CE** | RS.MI, RC.RP | Respond, Recover (RS, RC) | Respond | RA3000–RA5000 | Isolate, Evict, Restore | Technology | ICT continuity, backup |
| **AO** | RS.MI, DE.AE | Detect, Respond (DE, RS) | Respond | RA3000 | Isolate | Technology | Operational resilience |
| **FI** | RS.AN, ID.RA | Respond (RS) | Assess and decide | RA2000 | — | Technology, Process | Evidence, root cause |
| **RV** | ID.IM, PR.IR | Improvement (ID.IM) | Learn lessons | RA6000 | — | Process | Testing (DORA TLPT, NIS2 testing) |
| **RG** | GV, ID.IM, RS.CO | Improvement (ID.IM), Govern (GV) | Learn lessons | RA6000 | — | Business (Governance), Services | Reporting clocks (24h/72h/1m) |

Three corrections were made to this table at v1.0 against the current published standards:

- **ISO/IEC 27035-1:2023 has no "Recovery" phase.** Its five process phases are *Plan and prepare*, *Detect and report*, *Assess and decide*, *Respond* and *Learn lessons*. Recovery activity sits inside *Respond*. Earlier editions of this table named Recovery as a phase; it never existed as one.
- **SOC-CMM has no "Governance" domain.** Its five domains are Business, People, Process, Technology and Services; governance is an *aspect* within the Business domain. The table now names the domain and the aspect.
- **The SP 800-61r3 column is expressed in CSF 2.0 Functions**, because r3 is a CSF 2.0 Community Profile and has no lifecycle phases of its own to cite.

No mappings were added at v1.0. A crosswalk is only useful while every row can be checked against a published source, and enlarging it for the sake of a release is how it stops being checkable.

### 10.2 Positioning against adjacent models

| Model | Relationship |
|---|---|
| **SOC-CMM** | Assesses the SOC as a function. TIR-CMM assesses whether the organisation can act — including the parts of response that sit outside the SOC (authority, business acceptance, recovery). Complementary. |
| **NIST CSF 2.0** | Reporting layer above TIR-CMM. TIR-CMM supplies the evidence CSF Respond/Recover asks for. |
| **TID-CMM** | Sibling module. Bonded through R4 and the bridge domain (§11). |
| **Gartner CTEM** | Asks whether exposure is exploitable and observable. TIR-CMM asks whether exploitation can be stopped. |
| **DORA / NIS2** | TIR-CMM produces the testing, reporting-clock and continuity evidence these regimes require, without being a compliance model. |
| **VERIS** | Incident taxonomy; useful as an input to scenario derivation, not a maturity model. |

---

## 11. The TID-CMM bridge — resolving the IR domain overlap

TID-CMM v1.2 contains **IR — Incident Response & Recovery** at 10% with 6 sub-capabilities. Publishing TIR-CMM without addressing this leaves users with two contradictory response numbers.

**Resolution: substitution mode.** TID-CMM keeps IR at 10% — no reweighting, no invalidation of existing assessments — but the domain is redefined and gains a substitution rule.

### 11.1 Redefinition

TID-CMM's IR domain narrows from "incident response and recovery" to **the detection-to-response interface**: the handoff, not the response.

| Old scope | New scope (bridge) |
|---|---|
| Full IR capability, shallowly | Alert-to-case fidelity and enrichment |
| | Triage handoff quality and completeness of context |
| | Escalation path integrity from detection to responder |
| | Response feedback loop into detection tuning |
| | Case data sufficiency for tempo measurement |
| | Declared response maturity (or imported TIR-CMM score) |

This is genuinely within detection's remit — a detection function is accountable for whether its output is actionable — and it removes the pretence that a detection model can assess containment.

### 11.2 Substitution rule

```
if TIR-CMM assessment imported:
      TID-CMM.IR := TIR-CMM.overall
      report annotated "IR domain sourced from TIR-CMM v1.x, assessed <date>"
else:
      TID-CMM.IR := bridge sub-capabilities only
      report annotated "response capability not independently assessed"
```

Reciprocally, TIR-CMM consumes TID-CMM's constraint-adjusted overall as `D` in constraint R4. Because TID-CMM's IR is *substituted by* TIR-CMM rather than added to it, and because **R4 uses the pre-substitution TID-CMM score** (§8.4.1), there is no circular inflation. This ordering is normative in both specifications and the import contract carries the pre-substitution figure as a distinct field so that a consumer cannot get it wrong by accident.

### 11.3 Change required to TID-CMM

Version bump to **v1.3**, non-breaking:

- IR domain sub-capabilities rewritten (weight unchanged at 10%)
- Substitution rule added to the scoring page
- Export schema extended with `detection_score_pre_substitution` for R4 consumption
- Cross-link added: "Response assessed in depth by TIR-CMM"

Existing v1.2 assessments remain valid and comparable at overall level.

---

## 12. Assessment flow — ten steps

Mirrors TID-CMM's ten-step wizard so the two tools feel like one product. Target duration: **60–75 minutes** for a rapid baseline.

| # | Step | Input | Output |
|---|---|---|---|
| 1 | **Import or declare context** | TID-CMM JSON export, or manual entry | Crown jewels, attack paths, priority actors, detection score |
| 2 | **What you run** | Asset class selection (A1–A8) | Active columns of the lattice |
| 3 | **What you must keep running** | Crown jewel → asset class mapping, RTO/RPO | Cell criticality tiers T1–T3 |
| 4 | **Who targets you and how fast** | Priority actors, breakout times, and the source rank, reference and date of each | B(a) for tempo calculation, and its provenance for R5b |
| 5 | **Lattice scoping** | Auto-derived; assessor confirms N/A cells | In-scope lattice, 24–38 cells |
| 6 | **Containment options** | Per asset class: available primitives and tiers | Feeds CE domain, informs cell status |
| 7 | **Authority map** | Who may authorise what, when, with what latency | Feeds RA domain — the step that surprises people |
| 8 | **Lattice status** | RRS 0–3 per in-scope cell, **with the proof date and any recency triggers** | VRS, blind cells, expiries |
| 9 | **Capability scoring** | 58 sub-capabilities, 0–5, with evidence level **and evidence date** | Domain scores |
| 10 | **Tempo and results** | MTTD / MTTDecide / MTTC, with sample size, window and whether they are measured or estimated | Full report + ranked roadmap |

Step 7 is placed before step 8 deliberately. Assessors who map authority first score the lattice more honestly, because they have just discovered how many actions require an approval they cannot obtain at 03:00.

Steps 8, 9 and 10 each acquired a date or provenance field at v1.0. They are the fields §3.6 and R5b run on, and an assessment that skips them will grade down rather than fail.

### 12.1 Assessment modes

| Mode | Duration | Depth |
|---|---|---|
| **Rapid baseline** | Half a day | Self-assessment; identifies the conversations worth having |
| **Structured assessment** | Two weeks | Evidence-backed, stakeholder-participatory, R3 fully enforced |
| **Continuous** | Ongoing | Lattice status maintained live; validation expiry automated; reassessed quarterly |

---

## 13. Data model and interop contract

### 13.1 Principles inherited from TID-CMM

- **Entirely client-side.** No server calls, no analytics, no transmission. Non-negotiable — it is why practitioners will run it on real data.
- **Machine-readable model.** JSON definitions of domains, sub-capabilities, lattice, constraints, tempo intervals, breakout sources and recency rules published alongside the tool, and canonical over this document.
- **Free to use.** Model CC BY-ND 4.0; schemas and the machine-readable model CC BY 4.0; tooling source-available.
- **Offline parity.** Excel workbook with live formulas mirroring the tool exactly.

### 13.2 Import contract from TID-CMM

```json
{
  "schema": "tid-cmm/export/1.3",
  "assessed_at": "2026-08-16",
  "detection_score_pre_substitution": 2.34,
  "domains": { "TI": 2.6, "TM": 2.4, "DC": 2.1, "DE": 2.5,
               "AV": 1.9, "AA": 2.4, "IR": 2.2, "GV": 2.5 },
  "crown_jewels": [ { "id": "CJ-01", "name": "...", "asset_classes": ["A1","A7"] } ],
  "attack_paths": [ { "id": "AP-01", "actor": "...", "stages": ["S1","S3","S4","S7"],
                      "asset_classes": ["A2","A1","A3","A7"] } ],
  "actors": [ { "id": "TA-01", "name": "...", "breakout_minutes": 62 } ],
  "in_scope_techniques": [ { "id": "T1078", "status": 2, "mitigations": ["M1032"] } ]
}
```

`detection_score_pre_substitution` is the field constraint R4 consumes. `domains.DC` is the telemetry figure used for the C13 diagnostic (§14) and for nothing else.

### 13.3 Export contract

Every export states the model, tool and schema versions separately (see *Artifact versions*, front matter).

```json
{
  "schema": "tir-cmm/export/1.0",
  "model_version": "1.0",
  "tool_version": "1.0.0",
  "schema_version": "1.0",
  "tier": "baseline",
  "overall": 2.32,
  "self_assessed": 2.69,
  "band": "L1",
  "band_capped_by": ["R5", "R6"],
  "domains": { "RP": 2.52, "RA": 2.05, "RE": 2.52, "CE": 2.52,
               "AO": 2.52, "FI": 2.52, "RV": 1.52, "RG": 2.52 },
  "coverage": { "subs_scored": 58, "subs_not_applicable": 0, "subs_unanswered": 0,
                "domains_scored": 8, "weight_covered": 1.0 },
  "vrs_cj": 0.446,
  "engineered_rate": 0.406,
  "proven_rate": 0.0,
  "declared_proven_rate": 0.031,
  "blind_cells_t1t2": [ { "stage": "S3", "asset_class": "A1", "tier": "T1" } ],
  "tempo": [ { "actor": "TA-01", "breakout_min": 62,
               "breakout_source_rank": 4, "breakout_estimated": true,
               "timings_estimated": false, "sample_size": null,
               "mttd_min": 180, "mttdecide_min": 95, "mttc_min": 40,
               "containment_margin_min": -253, "tempo_ratio": 5.08 } ],
  "expiries": [ { "cell": "S2-A2", "from": 3, "to": 2,
                  "reason": "No date recorded for the last proof." } ],
  "constraint_adjustments": [
    { "constraint": "R1", "domains": ["RP","RE","CE","AO","FI","RG"], "ceiling": 2.52 }
  ],
  "roadmap": [ { "rank": 1, "kind": "capability", "id": "RV-2", "impact": 100.0 } ]
}
```

A domain nobody scored is exported as `null`, never as `0`. A consumer must be able to tell "unmeasured" from "measured and terrible".

`band_capped_by` may contain `R5b`. A consumer that enumerates constraint identifiers must accept it.

UTIOM's roadmap tool consumes both TID-CMM and TIR-CMM exports and produces the unified improvement plan across all three pillars.

---

## 14. Consistency checks

Thirteen checks run against a completed assessment. **They are informative, not normative. They never change a score.** They ask the assessor to look again.

### 14.1 Why they are not constraints

A constraint is a rule about what a score *can mean*. It fires deterministically, it lowers a result, and it cannot be argued with. A consistency check is a rule about what a set of answers *usually implies*. It fires on a pattern that is usually a scoring error and sometimes is not.

Turning these into constraints would be the wrong move twice over:

1. **It would double-count.** Most of these patterns already have a constraint governing the underlying weakness. Automation rated far above authority is already capped by R2; a positive containment margin on an estimated breakout is already capped by R5b. Adding a second arithmetic penalty for the same fact would punish it twice, which is precisely what §8.4.1 refuses to do with the telemetry domain.
2. **It would harden a heuristic into arithmetic.** "High recovery score with no timed restoration" is very often a scoring error. It is not always one — an organisation that has just completed a real restoration under incident conditions has the evidence and no exercise record. A constraint cannot make that distinction; a human reading a warning can.

So each check states four things: **which answers conflict**, **why the conflict matters**, **whether a constraint already governs it**, and **what to verify**. The fourth is the point. A check that only says "this looks wrong" wastes the assessor's time.

Checks carry a severity — critical, serious or warning — which orders them for reading. Severity has no arithmetic effect either.

### 14.2 The thirteen checks

| ID | Severity | Fires when | Why it matters | Already constrained by | Verify |
|---|---|---|---|---|---|
| **C1** | Critical | AO ≥ 3.5 and RA ≤ 2.0, as answered, before constraints | An automated containment action nobody may trigger without a meeting is a script, not automation. This is the commonest way a response programme reads as mature and performs as improvised. | R2, where it binds | That RA-1's pre-authorisation genuinely covers the actions AO-2 claims to automate, for the same asset classes, out of hours |
| **C2** | Critical | CE ≥ 3.5 with ≥ 8 in-scope cells, and ≥ 60% of them at status 0 or 1 | The domain score describes the capability you believe you have; the lattice records where it actually exists. A wide gap usually means the domain was scored on intent. | — | Walk three in-scope cells at status 0 or 1 and ask what would actually happen at 03:00, then re-score either the cells or the domain |
| **C3** | Critical | RV ≥ 3.0 with in-scope scenarios, none carrying a last-exercised date | Validation is the one domain that cannot be evidenced by intent. It either happened on a date or it did not, and every other domain is ceilinged by this one. | R1, where it binds | Record the last exercise date per scenario. With no date, the honest score is lower |
| **C4** | Serious | FI ≥ 3.5 while FI-3 (chain of custody) ≤ 1 | Forensic capability that has never preserved evidence to an admissible standard fails at the moment it is needed, and the failure is not recoverable afterwards. | — | That a preservation action has been executed and reviewed within twelve months, not merely documented |
| **C5** | Serious | CE ≥ 3.0 with in-scope scenarios and no observed recovery time recorded | An untimed recovery plan is a recovery intention. RTO commitments made without a measured restoration are the commonest cause of a missed recovery objective. | — | Time one restoration end to end, including integrity verification, and record it |
| **C6** | Serious | Containment margin ≥ 0 while tempo is estimated | A comfortable margin against a borrowed number is the most reassuring wrong answer the model can produce. The adversary you actually face sets the denominator. | R5b | Obtain an actor-specific figure, or measure breakout in the next purple-team exercise |
| **C7** | Warning | Tempo sample size below 10 timed incidents | With a handful of incidents the mean is dominated by whichever one went worst or best, and P90 cannot be estimated from four points. | — | Keep timing incidents; report P50 and P90 alongside the mean at ten or more |
| **C8** | Serious | RG ≥ 3.5 while GV-4 (independent calibration) is answered no | A governance score is a claim about how trustworthy the rest of the assessment is. Self-graded governance is circular, not merely optimistic. | R7, where it binds | Have a second assessor independently re-score everything rated above 3 and compare |
| **C9** | Critical | Overall ≥ 3.0 with no crown jewels declared | Threat-informed means response is shaped by what must not fail and how an adversary would reach it. Without a crown-jewel list every cell is weighted the same and the band name cannot mean what it says. | — | Agree a crown-jewel list with the business, then re-scope the lattice against the paths that reach it |
| **C10** | Warning | Three or more sub-capabilities scored ≥ 4 with VC3 evidence and no evidence date | VC3 means repeatable, current proof. Without a date, "current" cannot be checked, and validation that has quietly expired is indistinguishable from validation that holds. | §3.6.2 grades them to VC2 | Record the date of the most recent proof for each; anything older than twelve months should be re-evidenced or re-scored |
| **C11** | Serious | RE ≥ 3.5 and no asset class in the authority map names a decision-maker | A playbook without a named authority stops at the step that needs permission. That step is almost always the one that matters. | — | Name the individual or role entitled to authorise containment per asset class, and record their out-of-hours path |
| **C12** | Serious | Assurance depth with any of GV-3, GV-4, GV-5 unconfirmed | Assurance is the only depth with no band ceiling, and it earns that by being independently reviewable. Without a rubric, a second reviewer and a declared conflict-of-interest position, it is a long Baseline. | R7 caps at L4 | Complete the governance core, or record the assessment as Baseline depth |
| **C13** | Warning | TID-CMM telemetry domain sits ≥ 1.00 below overall detection maturity | Response depends on seeing the activity at all. A detection programme strong on content and thin on collection tends to fail on exactly the asset classes response most needs. | R4 applies one ceiling from overall detection maturity — this is context, not a second penalty | Which asset classes the telemetry gap falls on, and whether those are the classes carrying the Tier-1 lattice cells |

C13 is the check that closes §8.4's open question. It is the entire mechanism by which the telemetry figure influences an assessment: it raises a question for a human, and it does not touch the arithmetic.

### 14.3 What a consistency check is not

A check firing does not invalidate an assessment, does not require a re-score, and does not appear in `band_capped_by`. A check *not* firing is not a clean bill of health — thirteen patterns are thirteen patterns, not a complete account of the ways an assessment can be wrong.

---

## 15. Known limitations and open questions

### 15.1 Limitations of this release

Stated because a maturity model that will not describe its own maturity has failed its first test.

1. **No practitioner pilot has been run.** The model has been executed numerically against a fixture (§8.9) and reviewed against adjacent frameworks. It has not been run end to end by an assessor against a real estate, and the numerical validation is not a substitute for that. Every claim about how long an assessment takes is a design target, not an observation.

2. **The model has no external validation.** No independent party has assessed whether the domains are complete, whether the weights are defensible, or whether the bands describe recognisable populations. The weights are argued for in this document; they are not empirically derived, and no data exists against which they could be.

3. **Inter-assessor variance is unmeasured.** Two competent assessors scoring the same organisation may not produce the same result, and the size of that spread is unknown. The level-3 anchors, evidence levels and the §4.2.1 interval definitions all exist to narrow it, and none of them has been tested for whether they do. This is the single most important thing a pilot would measure.

4. **The OT/ICS profile is unpiloted.** Annex A is a design published as an extension specification. It has not been run against an OT estate, reviewed by an OT engineering authority, or tested for whether its questions are answerable by the people who would have to answer them.

5. **R1's calibration is settled by internal sensitivity testing only.** R1 binds six of eight domains in the worked example and subsumes R2 and R4 entirely. Alternatives were tested internally — a ceiling of RV + 1.5, and exempting RP from R1 — and RV + 1 was retained because the alternatives weakened the honesty mechanism more than they improved the diagnostic spread. That is an internal judgement against one fixture. It is not field evidence, and the independent-effect reporting rule mitigates the dominance rather than resolving it.

6. **Effort tiers are assigned, not measured.** The 20 / 23 / 15 spread across effort tiers is the model's judgement about what each item takes. No organisation has been timed doing any of it.

### 15.2 Open questions

Questions closed since v0.2 are not restated here. R4's basis is settled at §8.4.1; estimated tempo is settled and implemented as R5b; threat modelling without TID-CMM is settled by the prerequisite check.

1. **Lattice vs domain coupling.** The lattice constrains the band (R6) and drives prioritisation, but does not feed the overall arithmetic. An alternative is to make it a ninth pseudo-domain. The recommendation remains to keep them separate, matching TID-CMM's treatment of VCS, but the case has been argued rather than tested.

2. **Prerequisite calibration.** Whether the six-of-eight answer threshold and the L4/L3 ceilings in §8.4.2 are set at the right level. Field data comparing prerequisite proxies against real TID-CMM scores would settle it; no such data exists.

3. **MTTDecide instrumentation.** Many case systems cannot distinguish "alert validated" from "containment authorised" without process change. §4.2.1 defines the interval precisely enough to instrument; whether organisations can actually produce the timestamps is unknown, and if they systematically cannot, the model's central metric is unmeasurable in practice for a large part of its audience.

4. **Recency window length.** Twelve months is argued from the observation that response capability rests on people and authority, which change faster than detection logic. It is not derived from measured decay. Whether the right figure is nine months, twelve or eighteen — and whether it should vary by domain or asset class — is open.

5. **A8 depth in the core.** Annex A puts OT depth in an optional profile so the universal core is not enlarged with OT-only questions. Whether that is the right split, or whether a minimal OT-specific question belongs in the core lattice, cannot be answered until the profile has been piloted.

6. **Practitioner-facing worked example.** §8.9 is a validation trace. A published worked example with the full 58-item score sheet and lattice, matching TID-CMM's presentation, remains outstanding.

---

## 16. Release status

| # | Artefact | State at v1.0 |
|---|---|---|
| 1 | This specification | Published |
| 2 | Machine-readable model | Published — canonical over this document |
| 3 | Worked example with full constraint trace | Validation trace published (§8.9); practitioner-facing version outstanding |
| 4 | Assessment tool — single-file, client-side | Published |
| 5 | Excel workbook with live formulas | Published |
| 6 | White paper | Outstanding |
| 7 | Sites: model and tool | Published |
| 8 | TID-CMM v1.3 bridge update | Specified (§11); not yet released |
| 9 | UTIOM roadmap tool integration | Outstanding |
| 10 | Practitioner pilot | Not started (§15.1) |

---

## Annex A. OT/ICS profile

**Status: normative extension. NOT YET VALIDATED.**

This annex is a design. It is published as an extension specification so that it can be reviewed and criticised before it is used, and it has **not been piloted**. No OT estate has been assessed with it, no OT engineering authority has reviewed it, and none of its questions has been tested for whether the people who would have to answer them can. Treat every statement in this annex as a proposal with a version number, not as validated model content.

### A.1 Why this is a profile and not core content

Asset class **A8 — OT / ICS / IoT & Specialist** already carries OT in the core lattice, as a full column with the same eight stages, the same criticality tiers, the same RRS scale and the same recency rule as every other asset class. An organisation with an OT estate can assess it today without this annex, and the containment primitives the core names for A8 — segment, safe-state, manual fallback — are the correct primitives.

What the core cannot express is that OT changes the *meaning* of several core questions rather than adding new ones. "Can you isolate this asset?" is a capability question on A1 through A7 and a safety question on A8. "Is containment pre-authorised?" has one answer structure in IT and a different one where a plant manager holds an overriding authority that security cannot appeal.

The optional profile exists precisely so that the universal core is not enlarged with OT-only questions. Roughly nine in ten assessments will never scope A8 at all, and adding OT-specific sub-capabilities to the core would impose the cost on all of them, dilute every domain weight, and produce a model that is worse at its main job in order to be adequate at a specialist one. A profile that attaches to the A8 column, and only where A8 is in scope, keeps that cost where it belongs.

The profile does not create A8 coverage, does not change any domain weight, and does not alter the core scoring arithmetic. It refines how the A8 column's cells are scored and what evidence they require.

### A.2 Profile content

Eight areas. Each is expressed as a scoring qualifier on existing A8 lattice cells, not as a new domain.

#### A.2.1 Safety authority and who can overrule containment

In IT, security authority and operational authority can be reconciled by escalation. In OT they are frequently held by different functions with different legal obligations, and the safety authority wins by construction.

The profile requires the assessment to record, per A8 cell: who holds safety authority for the process the asset controls, whether that authority can veto a containment action, and whether the veto has a defined time limit. An A8 cell cannot be scored **Engineered** unless the answer to the second question is known in advance. A pre-authorisation that has never been tested against a safety veto is not pre-authorisation; it is an untested assumption about who is in charge.

#### A.2.2 Safe-state requirements

Containment on an OT asset is only permissible from, or into, a defined safe state. The profile requires the safe state to be named, its entry conditions documented, and the time to reach it recorded, per process rather than per device.

Where a safe state cannot be reached within the containment window, the honest RRS is **Manual only** at best, whatever the tooling can technically do. A containment action that would leave a process in an undefined state is not a containment option; it is an incident of its own.

#### A.2.3 Manual fallback

Manual fallback is the OT equivalent of a rollback path, and the profile treats it with the same requirement §6's CE-2 applies to blast radius: it must exist, be documented, and have been executed.

The assessment records whether manual operation is possible for the process, for how long it can be sustained, how many people it requires, and when it was last practised. Manual fallback that exists on paper and has never been run is scored as absent, on the same principle that an unrehearsed playbook is an assumed capability.

#### A.2.4 The inability to isolate or reimage

Several core containment primitives simply do not exist on A8. A PLC cannot be network-isolated without stopping the process it controls. A controller cannot be reimaged mid-cycle. An endpoint agent may be unsupported, prohibited by the vendor, or physically impossible on the hardware.

The profile requires these to be recorded as **structural absences of the primitive**, not as low scores against a primitive that could be built. The distinction matters for the roadmap: a missing IT primitive is a work item, and a missing OT primitive is a design constraint that the response plan must route around. An A8 cell whose only available action is "stop the process" is scored honestly as such, and the business impact of that action is recorded alongside it.

#### A.2.5 Engineering-owner approval paths

The people who must approve action on an OT asset are usually engineers, not security, and frequently not employees. The profile records the approval path per process: the named engineering owner, their out-of-hours availability, whether the approval is written or verbal, and the realistic latency of obtaining it at 03:00 on a Sunday.

This feeds the RA domain's decision-latency measurement, and it is where an OT assessment most often produces a number that surprises the organisation.

#### A.2.6 Operational continuity obligations

An OT process may carry a continuity obligation that is regulatory, contractual or physical — a chemical process that cannot be stopped safely at arbitrary points, a supply obligation with statutory penalties, a medical service that cannot be suspended.

The profile requires the obligation to be recorded per process and reconciled with the containment tiers. Where a continuity obligation forbids a containment tier outright, that tier is recorded as unavailable rather than as available-but-undesirable. A containment option nobody may lawfully exercise is not an option.

#### A.2.7 Vendor and integrator dependencies

Much OT response cannot be executed by the organisation at all. Vendor-supported systems, warranty conditions and integrator-held credentials mean the effective responder is a third party with a contractual response time.

The profile records, per A8 cell: whether execution requires a vendor or integrator, the contracted response time, whether it has ever been invoked, and what happens if it is not met. A cell whose response depends on an untested vendor SLA cannot be scored **Proven**, because the organisation has not proven it and neither has anybody else.

#### A.2.8 Recovery sequencing

OT recovery is ordered. Systems must be restored in a sequence determined by process dependencies, and restoring in the wrong order can be worse than not restoring at all.

The profile requires the recovery sequence to be documented per process, with dependencies named, and requires that the sequence has been walked through — at minimum on paper with engineering present. RTO for an OT process is measured from the start of the sequence to a verified safe operating state, not to the moment the last system booted.

### A.3 How the profile is applied

- It applies only where **A8 is in scope**. Where A8 is out of scope, the profile is inert and no question is asked.
- It **qualifies** A8 cell scoring; it adds no domain, no sub-capability and no weight.
- Its evidence requirements sit on top of the standard evidence levels. An A8 cell claiming VC3 must satisfy both §8.3 and the relevant profile requirement above.
- Profile findings appear in the report under the A8 column and in the roadmap as lattice items, ranked by the standard formula in §9.2.

### A.4 What would validate it

A pilot against at least one OT estate, answering three questions: whether the eight areas are the right eight, whether the questions are answerable by the engineering owners who must answer them, and whether the profile changes any A8 cell's score compared with assessing A8 under the core alone. If the answer to the third is no, the profile is documentation rather than a model extension, and should be published as guidance instead.

---

*TIR-CMM is a module of UTIOM. It measures whether an organisation can act on what it detects, inside the time the adversary allows, with the authority to do so — and whether it can prove it.*
