Specification §4, §7, §8, §9

Scoring & constraints

Two headline metrics, one arithmetic, seven constraints that only ever lower a score. The mechanism is deliberately mechanical: it cannot be argued with, and it is designed to make the gap between what an organisation believes it can do and what it has proven unavoidable rather than flattering.

TIR-CMM scores in the same shape as TID-CMM, so the two results sit on one axis and the workbook formulas transfer. What is new is on this page: a metric for the race against adversary tempo, a constraint that caps capability at the authority to use it, and a constraint that caps response at the detection it depends on.

Two things changed in v0.2. Constraint R3 no longer asks a yes/no question about evidence — it applies a graduated evidence cap derived from an evidence level of 0 to 3. And a seventh constraint, R7, caps the band at the depth of the assessment and the strength of its governance.

Two headline metrics

The domain scores answer how mature is the capability. These two answer the questions a board actually asks: where can we act, and are we faster than the adversary.

Validated Response Score (VRS)

The direct analogue of TID-CMM's Validated Coverage Score, computed across the in-scope cells of the Containment Lattice.

Unweighted:

VRS = Σ RRS(c) / (3 × N)          for all in-scope cells c, N = |in-scope cells|

Crown-jewel weighted — the reported figure:

VRS_cj = Σ (ω(c) × RRS(c)) / (3 × Σ ω(c))

Three further rates are reported alongside VRS, and they are expected to be the numbers that land in the room:

Engineered rate

The proportion of in-scope cells at status 2 or above — “we believe we can act here.”

Proven rate

The proportion of in-scope cells at status 3 — “we have shown we can act here.”

Blind cells

The count of status-0 cells at tier T1 or T2 — the places on a path to a crown jewel where you have no option at all.

The expected finding, mirroring TID-CMM's 48.9% claimed against 12.7% proven, is a large gap between the engineered rate and the proven rate. The model is built to make that gap unavoidable rather than to flatter it.

Containment Margin (CM)

Detection maturity is a coverage problem. Response maturity is a race. A model that scores response without reference to adversary tempo is measuring paperwork. For each priority actor a with breakout time B(a):

CM(a) = B(a) − ( MTTD + MTTDecide + MTTC )
TR(a) = ( MTTD + MTTDecide + MTTC ) / B(a)          "Tempo Ratio"
The four terms of the Containment Margin, and where each is sourced.
Term Definition Source
B(a) Breakout time for the actor — time from initial foothold to first lateral movement Threat intel; UTIOM metrics calculator; industry baseline where unavailable
MTTD Mean time to detect Imported from TID-CMM / SIEM
MTTDecide Mean time from validated alert to authorised containment decision TIR-CMM — measured here for the first time
MTTC Mean time to execute containment once authorised TIR-CMM

MTTDecide is the contribution. Every maturity model in circulation folds decision latency into MTTR, where it disappears. Separating it exposes the most common and most fixable failure in incident response: the organisation is not slow at containing, it is slow at being allowed to contain. Organisations that measure it routinely discover MTTDecide exceeds MTTC by an order of magnitude.

Tempo Ratio interpretation. The ratio drives constraint R5.
Tempo Ratio Meaning
TR < 0.5 Containment lands well inside breakout — genuine tempo advantage
0.5 ≤ TR < 1.0 Contains before spread, with little margin
1.0 ≤ TR < 2.0 Structurally behind — you contain a spread intrusion, not a foothold
TR ≥ 2.0 You are performing post-incident recovery and calling it response

Tempo Ratio drives constraint R5.

Scoring mathematics

Identical in form to TID-CMM, so results are comparable and the workbook formulas transfer. Every sub-capability is scored 0–5; sub-weights are within-domain and total 100 per domain; domain weights total 100.

Sub-capability to domain

domain_score = Σ(sub_weight × sub_score) / Σ(sub_weight)

Not-applicable sub-capabilities are removed from both numerator and denominator. No penalty, no benefit — the same treatment the lattice gives an out-of-scope cell.

Domain to overall

overall_score = Σ(domain_weight × domain_score) / Σ(domain_weight)

Why the lattice is reported alongside, not folded in

The lattice does not feed the overall domain score arithmetically. It is reported alongside it, exactly as TID-CMM reports Validated Coverage Score alongside domain maturity. Conflating them would let broad shallow coverage mask a structural inability to act.

The lattice does, however, bind the score in two other ways: through constraint R6, which caps the band when a Tier-1 cell has no response option at all, and through prioritisation, where cell gaps compete directly with capability gaps for the top of the roadmap.

Reported output set

What a completed TIR-CMM assessment reports.
Output Form
Overall maturity0.00–5.00 + band L0–L5
Domain scores8 × 0.00–5.00, radar chart
VRS_cj0–100%
Engineered rate / Proven rate0–100% each — the headline gap
Blind cells (T1/T2)Count + named list
Containment Margin per priority actorMinutes, signed
Tempo Ratio per priority actorRatio
Constraint adjustmentsSelf-assessed vs adjusted, per constraint
Ranked roadmapOrdered list with impact scores

Integrity constraints

TID-CMM's central innovation is that constraints apply mechanically, lower scores only, and cannot be argued with. TIR-CMM inherits this exactly and adds three constraints of its own: an authority ceiling, a tempo ceiling, and — new in v0.2 — an assessment-depth and governance ceiling.

Application order

Each constraint operates on figures already adjusted by the preceding ones.
# Step Operates on
1R3 — Evidence capSub-capability level
2Compute domain scoresSub-capabilities → domains
3R4 — Detection dependencyDomain level, cross-model
4R2 — Authority ceilingDomain level
5R1 — Rehearsal ceilingDomain level
6Compute overall scoreDomains → overall
7R5 — Tempo cap · R6 — Blind-cell cap · R7 — Depth & governance capBand only, applied together; the lowest cap wins

Reporting rule — independent effect, not marginal effect. Because R1 is applied last and is usually the tightest ceiling, it routinely subsumes R2 and R4, which would then appear in the report as “no effect” even where they are diagnostically the most important finding. The tool must therefore compute and display each constraint's independent effect — what it would cap, evaluated against the pre-constraint domain scores — alongside its marginal effect on the final number. Without this, R2 and R4 become invisible in exactly the organisations they were written for. This behaviour was identified during numerical validation of the worked example.

R1 — Rehearsal Ceiling

∀ domain d ≠ RV:   d ≤ RV + 1

An unrehearsed playbook is an assumed capability.

The direct analogue of TID-CMM's C1. Response is more exposed to this than detection: a detection rule at least executes automatically every day, whereas a playbook that is never run may not survive first contact with the tooling, the people or the clock.

R2 — Authority Ceiling Unique to TIR-CMM

AO ≤ RA + 1
CE ≤ RA + 1

Automation you are not permitted to fire is a demonstration, not a capability.

The constraint with no equivalent in any existing model. It expresses the single most common structural failure in incident response: a technically capable organisation that cannot act because nobody with authority is available, informed, or willing to accept the business impact. Buying a better SOAR does not move this score.

R3 — Evidence Cap Rewritten in v0.2

effective_sub_score = min( sub_score, EVIDENCE_CAP[ evidence_level ] )

EVIDENCE_CAP = { 0: 1,  1: 2,  2: 4,  3: 5 }

Unsupported claims must not outrank evidenced assessments.

v0.1 asked a yes/no question — is there a named, dated artifact? — and demoted any 4 or 5 that failed it to 3. That was a blunt instrument in both directions. It let a policy document carry a score of 3 unchallenged, and it treated a repeatably validated capability and a screenshot as the same answer.

v0.2 replaces the flag with an evidence level of 0 to 3, and each level caps the score it will support. The levels are aligned to TID-CMM's Validated Coverage grades, so an imported VC grade sets the ceiling directly instead of having to be reinterpreted.

Defined in assets/model.js as EVIDENCE and EVIDENCE_CAP. Full descriptions on the model page.
Level VC grade Name What it expects Caps at
0 VC0 Assertion only None, or a verbal claim. 1
1 VC1 Design or policy A document, policy, screenshot, or partial implementation. 2
2 VC2 Implemented and tested Implementation evidence plus a test result, ticket, or system record. 4
3 VC3 Repeatably validated Recent, repeatable, independently reviewable proof from the live environment. 5

The cap is a ceiling, not a conversion. A high evidence grade permits a high score; it does not create one. VC3 evidence against a capability the assessor scores 2 leaves the score at 2. The cap only ever removes the part of a claim its evidence cannot carry — it never adds.

A blank evidence level is treated as level 0 — assertion only — capping that sub-capability at 1. It is not skipped and it is not waved through. An assessment that declines to say what its evidence is has answered the question.

The one exception is Pulse, which does not ask for evidence at all. Rather than penalise twenty answers as though evidence were absent, the cap is switched off for Pulse and the reduced rigour is carried by the L3 band ceiling under R7.

v0.1 records still resolve without rescoring: the old boolean flag meant a named, dated artifact, which is VC2, so true imports as level 2 and caps at 4.

The graduated cap is materially stricter than the flag it replaces. In the worked example below it binds six sub-capabilities where the v0.1 rule bound three, because a level-1 answer now caps at 2 rather than passing at 3 — and every one of those six is a capability the organisation believed it had.

R4 — Detection Dependency Cross-model hinge

RE ≤ D + 1
CE ≤ D + 1
FI ≤ D + 1

where D = TID-CMM overall score, if a TID-CMM assessment is imported
        = assessor-declared detection maturity (0–5), otherwise, flagged "unverified"

You cannot respond to what you never saw.

This is the structural hinge that makes TIR-CMM complete TID-CMM rather than duplicate it. It also creates the correct incentive: an organisation cannot reach a high response score by buying response tooling while remaining blind.

The unverified path is deliberately uncomfortable — the report states plainly that the response score rests on an unaudited detection claim, and caps the overall band at L3.

R4 also subsumes the role of TID-CMM's C4 (intent ceiling) for the response module, since the threat modelling that produces the lattice is imported from TID-CMM's TM domain rather than re-assessed here.

R5 — Tempo Ceiling

Applies to the band, not to domain scores.

TR = (MTTD + MTTDecide + MTTC) / B(a_priority)

TR ≥ 2.0            ⟹  band capped at L1
1.0 ≤ TR < 2.0      ⟹  band capped at L2
tempo not supplied  ⟹  band capped at L3, report marked "tempo unverified"

A response capability structurally slower than the adversary is not a level 3 capability, whatever the documentation says.

This is the constraint that prevents TIR-CMM from becoming a paperwork audit. An organisation can hold excellent playbooks, full automation coverage and a rehearsal programme, and still be losing every race. The tempo cap makes that visible in one number that a board understands.

R6 — Blind-Cell Gate

∃ cell c : tier(c) = T1 ∧ RRS(c) = 0   ⟹  band capped at L2

If there is a stage on a crown jewel where you have no response option at all, you are not threat-informed.

The response analogue of TID-CMM's structural-gap analysis, and consistent with UTIOM's Law 3 — crown jewels drive resource allocation — and Law 5 — blind spots are architectural choices.

R7 — Assessment Depth & Governance Ceiling New in v0.2

Applies to the band, not to domain scores.

tier = Pulse                                        ⟹  band capped at L3
tier = Baseline                                     ⟹  band capped at L4
self-assessment: ¬(evidence-led scoring
                 ∧ independent calibration
                 ∧ assessor ≠ approver)             ⟹  band capped at L4

Twenty questions cannot evidence adaptive maturity, and a result nobody challenged is not assurance.

R7 does two jobs with one mechanism. It records how much of the organisation the assessment actually reached, and it records whether anybody independent challenged what it found.

The depth half follows from the tiering. A Pulse assessment scores 20 of the 58 sub-capabilities, has no lattice and asks for no evidence, so it caps at L3. A Baseline assessment scores all 58 with evidence levels and a full lattice, but is self-service, so it caps at L4. Assurance carries no tier cap.

The governance half is independent of tier and applies to any assessment. Three of the eight assurance conditions decide it: GV-3 evidence-led rather than self-declared scoring, GV-4 independent calibration of every score above 3, and GV-5 separation of the person proposing scores from the person approving them. Miss any one of the three and the result is a self-assessment, capped at L4.

The two halves compose, and the lowest cap wins. Commissioning an Assurance engagement does not lift the governance cap on its own — running it with an independent validator does.

The ceiling is not a penalty. The arithmetic underneath is identical at every tier; a Pulse assessment computing 3.4 still reports 3.4. What the ceiling records is that the depth an assessment reaches is itself evidence about how far its result can be trusted. A self-assessment is useful — it is often exactly the right thing to run — but it must not be presented as assurance, and R7 is what stops the report doing so quietly.

Worked example — “Meridian Group”

The full pipeline was executed numerically against a fictional mid-size European financial services organisation of around 9,000 staff, with a competent but optimistic security team: good preparation, real detection engineering investment, a well-regarded SOAR deployment, tabletop exercises — and no live-fire, no measured decision latency, nothing signed on containment authority. This is the most common profile in the market.

Structural checks passed: domain weights sum to 100, 58 sub-capabilities, and all eight domains' sub-weights sum to 100.

Self-assessed
2.69

Would have read L2 — “approaching level 3” in most models

Adjusted
2.32

After R3, R4, R2 and R1 — Δ −0.36

Proven rate
3.1%

Against an engineered rate of 40.6%

Tempo Ratio
5.08

Containment Margin −253 min

Reported band
L1

Documented — capped by R5, with R6 and R7 also firing

Constraint trace

Each constraint applied in order, against figures already adjusted by its predecessors.
Step Effect
Self-assessed overall 2.69 (L2)
R3 evidence cap 6 sub-capabilities capped by their evidence level, all at level 1 — design or policy only, so capped at 2: RP-3 4→2, RP-4 4→2, RP-5 3→2, RA-1 4→2, RE-5 3→2, AO-5 3→2. Under the v0.1 yes/no flag only three moved, and only from 4 to 3.
R4 detection dependency Ceiling 3.34 — not binding at final ordering; independently, would not bind
R2 authority ceiling Ceiling 3.05 — not binding at final ordering; independently, would not bind
R1 rehearsal ceiling RV = 1.52 → ceiling 2.52. Binds six of eight domains: RP 3.20→2.52, FI 2.85→2.52, AO 2.80→2.52, RG 2.72→2.52, CE 2.68→2.52, RE 2.66→2.52
Constraint-adjusted overall 2.32 (L2), Δ −0.36

Lattice

32 in-scope cells — inside the 24–38 design band, and four more than in v0.1 because the S0 Prevent row is now scored. VRS_cj 45.5%, engineered rate 40.6%, proven rate 3.1%. One blind Tier-1 cell: S3 × A1 — privilege escalation on identity infrastructure, no response option at all.

The engineered-versus-proven gap — 40.6% against 3.1% — is the response analogue of TID-CMM's 48.9% against 12.7%, and is proportionally wider. That is the expected and intended result, since response capability is rehearsed far less often than detection logic is exercised.

Tempo

Containment Margin for the priority actor.
Term Minutes
Breakout time B(a)62
MTTD180
MTTDecide95
MTTC40
Total time to contain315
Containment Margin−253
Tempo Ratio5.08

Band outcome

R5 fires — TR ≥ 2.0, cap L1. R6 fires — Tier-1 blind cell, cap L2. R7 fires — this is a Baseline, self-scored assessment with no independent calibration, cap L4. The lowest cap wins. Final reported band: L1 — Documented, against a self-assessed 2.69 that would have read L2 and, in most maturity models, been reported as “approaching level 3.”

This is the intended shape of the mechanism. TID-CMM's published worked example moves an organisation from a self-assessed 2.44 to an adjusted 2.34 — a deliberately modest drop that demonstrates the mechanism without theatre. TIR-CMM behaves comparably on the decimal, and the honest shock arrives in the band, not the number.

Note what R7 is doing here, and what it is not. It is not the binding cap — R5 is, by three bands — so it costs this organisation nothing. Its job is to be recorded in the report anyway, so that a reader knows the 2.32 was self-scored and never independently challenged. On an organisation whose tempo and lattice were healthy, R7 would be the cap that mattered.

Mechanical roadmap — top five

1 RV-2 — technical live-fire exercising
2 S3 × A1 — the blind identity cell
3 S0 × A1 — prevention and hardening on identity infrastructure
4 S0 × A7 — prevention and hardening on data stores and backup
5 RA-2 — decision-latency measurement (MTTDecide)

Not one of the top five is a product purchase, and three of five are things this organisation believed it already had. That is the model working.

The two S0 cells at ranks 3 and 4 are new in v0.2, and they arrive there mechanically rather than by advocacy: S0 carries the highest stage leverage in the model, and both cells sit on crown-jewel asset classes at status 1. The model is saying, without being asked to, that hardening identity and backup would buy this organisation more than any further response engineering — which is the correct answer for an organisation losing the race by a factor of five.

Two model defects were found and corrected by running this arithmetic: the roadmap scale mismatch, resolved by the mandatory normalisation rule below, and the constraint-reporting problem, resolved by the independent-effect reporting rule above. Both would have shipped unnoticed without executing the arithmetic.

Prioritisation and roadmap generation

Two ranked lists are produced and then merged. Both are mechanical — no advocacy, no facilitator bias.

Capability gaps — sub-capability level

Identical to TID-CMM's formula:

impact = (domain_weight / 100) × (sub_weight / 100) × gap_to_target × 1000

Lattice gaps — cell level

cell_impact = ω(c) × λ(stage) × (3 − RRS(c)) × 100

  ω = criticality tier weight (T1=3, T2=2, T3=1)
  λ = stage leverage (S0 1.6 · S1 1.5 … S7 0.8)

This surfaces the correct counter-intuitive result: a status-0 cell at S1 on a Tier-0 identity asset outranks a status-2 cell at S7 on a peripheral system, even though the latter feels more urgent during an incident.

From v0.2 the leverage scale starts at S0 — Prevent & Harden, λ = 1.6, the highest value in the model. A preventive gap on a crown jewel therefore outranks every response gap of equal status and tier, because prevention buys the one thing response cannot manufacture, which is time.

Normalisation before merge — mandatory

The two lists operate on incompatible scales. Maximum capability impact is 0.16 × 0.25 × 3 × 1000 = 120; maximum cell impact is 3 × 1.6 × 3 × 100 = 1440.

Merging the raw figures produces a roadmap consisting entirely of lattice cells, with every capability gap buried — verified numerically during model validation. Normalisation is therefore not optional.

Each list is normalised to its own maximum before merging:

normalised_impact = 100 × impact / max(impact in that list)

This is scale-free and survives any future reweighting. The merged list is then sorted descending. Validation confirms it produces a properly interleaved result whose top five are the live-fire exercising gap, the blind identity-escalation cell, the two crown-jewel prevention cells, and decision-latency measurement — which is the model's central argument, generated mechanically rather than asserted.

Merge and sequencing

  1. Anything that lifts a binding constraint is promoted, because it unlocks score elsewhere. In practice, RA and RV work is almost always promoted — which is the model's central argument made arithmetically.
  2. Items are grouped into a 90-day / 180-day / 12-month plan, matching UTIOM's roadmap tool.
  3. The output names capabilities and authorities required, not products — inherited directly from TID-CMM's design principle.