v0.2 — the headline change
Three tiers
One model, three depths. Twenty minutes, an afternoon, or a facilitated engagement. Each tier is a strict superset of the one below, so you can start in twenty minutes and deepen later without ever rescoring anything you have already answered.
The v0.1 model was too heavy to start with. It asked for 58 sub-capabilities, a containment lattice, an authority map and three tempo measurements before it would report anything at all — which is the right depth for a result you intend to defend, and entirely the wrong depth for the first conversation. Organisations that could not clear that bar did not get a shallower answer; they got no answer. That was a design failure, not a filter.
v0.2 fixes it by splitting the same model into three depths rather than diluting it into one. Nothing was removed. What changed is where you are allowed to stop.
The three tiers compared
| Tier | Time | What it covers | Band ceiling | Who it is for |
|---|---|---|---|---|
| Pulse | 20 minutes | Twenty questions, no lattice, no ATT&CK, no evidence review. Produces a band, the three things holding you back, and a quick-fix list. Indicative, not defensible — and it says so on the report. | L3 | A CISO with a spare coffee break, or a first conversation with a new client. |
| Baseline | 1–2 hours | All 58 sub-capabilities with evidence levels, the containment lattice scoped to your crown jewels, the authority map, and response tempo. Self-service and defensible enough to take to a leadership team. | L4 | A SOC or IR lead running an honest internal baseline. |
| Assurance | 2–4 weeks | Everything in Baseline plus telemetry attributes per asset class, exercised scenarios with observed timings, the evidence register, and the governance layer — decision rights, independent calibration and second review. Evidence-led interviews, not self-scoring. | None | A facilitated engagement with an assessment lead, business owners and an independent validator. |
Pulse
Where do we stand, roughly, and what should we do first?
Four steps: context, twenty questions, tempo, results.
Baseline
What is our response capability, and can we prove it?
Eleven steps: context, assets, crown jewels, actors, lattice scoping, containment, authority, lattice status, capability scoring, tempo, results.
Assurance
Can we stand behind this in front of a regulator or a board?
Fourteen steps: the Baseline eleven plus telemetry, scenarios and governance.
The ceiling is not a punishment
Each tier carries its own band ceiling: Pulse caps the reported band at L3, Baseline at L4, and Assurance carries no tier ceiling at all. That looks like a penalty for choosing the fast option. It is not.
The depth an assessment reaches is itself evidence about how far the result can be trusted. Twenty questions cannot evidence a claim of adaptive response maturity, however good the answers are — not because the answers are wrong, but because twenty answers do not carry enough of the organisation to support that claim. The ceiling records the strength of the method alongside the strength of the capability, which is exactly what a reader of the report needs in order to know what to do with it.
This is the same logic the model applies everywhere else. Constraint R3 caps a score at the strength of its evidence. Constraint R5 caps the band when tempo is unmeasured. Constraint R4 caps response at the detection it depends on. R7 extends it one level up: it caps the band at the strength of the assessment.
A ceiling does not lower your score. It states how much of the scale your method is entitled to reach.
The arithmetic underneath is identical in all three tiers. A Pulse assessment computing 3.4 reports 3.4 — it simply reports the band as L3 rather than L3 with an implied claim to more, and the report says which cap is doing the work. Deepening the assessment does not change the arithmetic; it removes the cap.
Pulse does not ask for evidence levels at all. That is deliberate. Twenty minutes is not enough time to review artifacts honestly, and asking for evidence you have not looked at produces worse data than not asking. Rather than penalise every Pulse answer as though evidence were absent, the model switches the R3 evidence cap off for Pulse and carries the reduced rigour in the L3 band ceiling instead. One mechanism, applied once, in the place where it is honest.
The superset guarantee
Each tier is a strict superset of the one below. Every question in Pulse is a question in Baseline, asked in the same words, scored on the same 0–5 scale, against the same anchors. Every question in Baseline is a question in Assurance. Nothing is reworded, rescaled or reinterpreted between tiers.
The practical consequence is the point of the design: deepening an assessment never costs you the work you have already done. Run Pulse on Monday, and the twenty answers you gave are already 20 of the 58 Baseline answers when you come back to it. Run Baseline, and every one of its eleven steps is already complete when you commission an Assurance engagement. There is no migration, no remapping, and no moment where an organisation is told its earlier assessment does not count.
This matters more than it sounds. The most common reason a maturity assessment is never repeated is that repeating it means starting again. A model that only rewards the deepest tier will be run once, at the wrong depth, by an organisation that was not ready for it — and then never again.
What each tier adds over the one below
| Step | Pulse | Baseline adds | Assurance adds |
|---|---|---|---|
| Context | Crown jewels, priority actors, detection figure | — | — |
| Capability scoring | 20 sub-capabilities, no evidence level | All 58, each with an evidence level 0–3 | Evidence-led interview rather than self-declaration |
| Estate and scope | — | Asset classes, crown-jewel mapping, actor breakout times, lattice scoping | — |
| Containment Lattice | — | 8 stages × 8 asset classes, containment options, authority map, cell status 0–3 | — |
| Tempo | MTTD, MTTDecide, MTTC | — | — |
| Telemetry | — | — | 8 telemetry attributes scored per asset class |
| Scenarios | — | — | 12 starter scenarios × 10 lifecycle stages, with observed timings |
| Governance | — | — | 8 assurance conditions, 10 calibration questions, decision rights, RACI |
Read down the Pulse column and you have the shortest honest assessment the model can produce. Read across a row and you have the reason the deeper tier exists.
Pulse — twenty minutes
Twenty sub-capabilities, chosen so that every one of the eight domains is represented and the selection is weighted towards the things that most often decide the outcome: authority, rehearsal, and whether containment options exist at all.
| Domain | In Pulse | Which ones, and why |
|---|---|---|
| RP — Response Preparation | 2 | RP-1 a plan exists; RP-3 someone answers at 03:00. |
| RA — Response Authority | 3 | RA-1 pre-authorisation; RA-2 decision latency; RA-4 out-of-hours authority. |
| RE — Response Engineering | 3 | RE-1 playbook coverage; RE-2 threat-informed; RE-3 playbooks as code. |
| CE — Containment & Eradication | 4 | CE-1 tiered options; CE-3 identity containment; CE-5 eradication; CE-7 backups. |
| AO — Automation & Orchestration | 2 | AO-2 automated containment; AO-6 case data completeness. |
| FI — Forensics & Investigation | 1 | FI-1 triage and scoping. |
| RV — Response Validation | 2 | RV-2 live-fire exercising; RV-4 timing captured under exercise. |
| RG — Response Governance | 3 | RG-1 metrics; RG-2 post-incident review; RG-3 feedback into detection. |
Domain scores in Pulse are extrapolated from whichever of that domain's sub-capabilities are in the subset. That is a real approximation, and the report states it rather than hiding it. FI is represented by a single question and its domain score should be read accordingly.
Pulse still runs the domain-level constraints. R1 still caps every domain at rehearsal plus one; R2 still caps automation and containment at authority plus one; R5 still caps the band when tempo is losing or unsupplied. Those are the constraints that produce the finding, and the finding is the reason to run Pulse at all.
Pulse is indicative, not defensible. It is designed to identify the conversations worth having, not to be quoted in a board paper as an assurance position. The report carries that statement, and the L3 ceiling enforces it.
Baseline — one to two hours
The model as v0.1 defined it, plus evidence levels: all 58 sub-capabilities, the Containment Lattice scoped to your crown jewels, the containment inventory, the authority map and the three tempo measurements.
Baseline is where the model does its characteristic work. It is the tier at which the lattice exists, so it is the first tier that can report the gap between engineered and proven coverage — the difference between believing you can act and having shown it. It is also the first tier that asks for an evidence level against every score, which is what makes constraint R3 bite.
The band ceiling is L4. A self-service assessment, however carefully done, is a self-assessment: nobody independent challenged the scores, and no second reviewer looked at the claims above 3. That is enough to support a serious internal improvement programme. It is not enough to support a claim of validated maturity.
Assurance — two to four weeks
Everything in Baseline, plus the three things that turn a populated questionnaire into an assurance result.
Telemetry attributes
Eight attributes, per asset class
Collection, completeness, timeliness, integrity and normalisation, retention, queryability, health monitoring and access protection — scored against each asset class rather than against each ATT&CK data component.
Scenario validation
Twelve scenarios, ten stages each
One business-relevant path from first evidence to trusted recovery, exercised and timed. Questionnaires reveal claims; scenarios reveal integration.
Governance layer
Eight assurance conditions
Named sponsor, frozen scope, evidence-led scoring, independent calibration, separation of assessor and approver, expiring risk acceptances, retested closure, defined reassessment triggers.
Assurance carries no tier ceiling — but it does not follow that an Assurance engagement automatically reaches L5. The other constraints all still apply, and the governance layer adds one of its own: without evidence-led scoring, independent calibration and separation of assessor from approver, the result is a self-assessment and is capped at L4 whatever tier it was run at. Buying the deeper engagement does not lift the ceiling. Doing the governance work does.
Each of the three additions has its own page section: telemetry attributes, scenario validation and the governance layer on the model page.
Which one to run
The honest default is start with Pulse. Twenty minutes costs nothing, the superset guarantee means the answers are not wasted, and the finding it produces is usually enough to decide whether the deeper tier is worth commissioning.
- You have twenty minutes and no prior assessment. Pulse. It will tell you which three things are holding you back, and that is a better starting position than a full assessment you never finish.
- You are the SOC or IR lead and you want an honest internal picture. Baseline. Block out an afternoon, bring the crown-jewel list and the case-system timings, and expect the authority map to surprise you.
- Someone external will read the result. Assurance. A regulator, a board, an insurer or an acquirer will ask who challenged the scores, and a self-assessment has no answer to that question.
- You already ran Baseline last year. Deepen rather than repeat. Response maturity is perishable — lattice cells expire, owners leave, platforms change — so the useful next assessment is the one that adds scenarios and independent calibration, not the one that re-answers the same 58 questions.