Assessment guide
How to run an assessment
Pick a depth, then work through its steps in order. Pulse takes twenty minutes, Baseline takes an afternoon, Assurance takes two to four weeks. The assessment is worth exactly as much as the honesty you bring to it, so most of this page is about how to be honest.
Whichever tier you run, the flow is deliberately ordered. Context and scope come first, because the lattice cannot be sized until the estate and the attack paths are declared. Authority comes before lattice status, because assessors who have just mapped their approval paths score containment readiness differently. Tempo comes last, because it is the number that decides the band.
Each tier is a strict superset of the one below, so starting shallow costs you nothing. Nothing you answer in Pulse has to be re-answered in Baseline, and nothing you answer in Baseline has to be re-answered in Assurance.
Which tier to pick
The honest default is start with Pulse. Twenty minutes costs nothing, the answers carry forward, and the finding it produces is usually enough to decide whether the deeper tier is worth commissioning.
| Tier | Time | Steps | Band ceiling | Pick it when |
|---|---|---|---|---|
| Pulse | 20 minutes | 4 | L3 | You have a coffee break, no prior assessment, and you want to know which three things to fix first. Also the right opening move with a new client. |
| Baseline | 1–2 hours | 11 | L4 | You are the SOC or IR lead and you want an honest internal picture you can take to a leadership team. |
| Assurance | 2–4 weeks | 14 | None | Somebody external will read the result — a regulator, a board, an insurer, an acquirer — and will ask who challenged the scores. |
Two things are worth knowing before you choose.
The ceiling is not a penalty. The arithmetic is identical at every tier; a Pulse assessment computing 3.4 reports 3.4. What the ceiling records is that the depth an assessment reaches is itself evidence about how far its result can be trusted. Deepening the assessment does not change the arithmetic — it removes the cap.
Assurance has no tier ceiling, but it does not automatically reach L5. Under R7, a result produced without evidence-led scoring, independent calibration and separation of assessor from approver is a self-assessment and caps at L4 whatever tier it was run at. Commissioning the engagement does not lift that. Doing the governance work does.
What each tier involves
Pulse · 20 minutes · caps at L3
Where do we stand, roughly, and what should we do first?
Twenty questions, no lattice, no ATT&CK, no evidence review. Produces a band, the three things holding you back, and a quick-fix list. Indicative, not defensible — and it says so on the report.
- Context. Crown jewels, priority actors, the detection figure.
- Twenty questions. Two to four sub-capabilities per domain, weighted towards authority, rehearsal and whether containment options exist at all.
- Tempo. MTTD, MTTDecide, MTTC.
- Results. Band, top three findings, quick fixes.
Domain scores are extrapolated from whichever sub-capabilities of that domain are in the subset. That is a real approximation and the report states it.
Baseline · 1–2 hours · caps at L4
What is our response capability, and can we prove it?
All 58 sub-capabilities with evidence levels, the containment lattice scoped to your crown jewels, the authority map, and response tempo. Self-service and defensible enough to take to a leadership team.
- Context — import or declare.
- What you run — asset classes A1–A8.
- What you must keep running — crown jewels, RTO, RPO.
- Who targets you — actors and breakout times.
- Lattice scoping — confirm the N/A cells.
- Containment options — primitives per asset class.
- Authority map — who may authorise what, and how fast.
- Lattice status — RRS 0–3 per in-scope cell.
- Capability scoring — 58 subs, each with an evidence level.
- Tempo — MTTD, MTTDecide, MTTC.
- Results — full report and ranked roadmap.
Assurance · 2–4 weeks · no ceiling
Can we stand behind this in front of a regulator or a board?
Everything in Baseline plus telemetry attributes per asset class, exercised scenarios with observed timings, the evidence register, and the governance layer. Evidence-led interviews, not self-scoring.
The Baseline eleven, with three inserted:
- Telemetry — after the authority map. Eight attributes scored against each asset class.
- Scenarios — after capability scoring. Twelve starter scenarios, ten lifecycle stages each, exercised and timed.
- Governance — after tempo. Eight assurance conditions, ten calibration questions, decision rights and RACI.
Details of all three are on the model page.
You do not need TID-CMM to use this
The tool runs standalone by default. A TID-CMM assessment makes the result stronger, but its absence is not a reason to be turned away — it is a reason to be explicit about what the result can and cannot support.
Constraint R4 — you cannot respond to what you never saw still applies exactly as written: detection maturity caps RE, CE and FI. What has changed is that the detection figure R4 uses can come from three places, and the model treats them differently, because they carry different amounts of proof.
| Where the detection figure comes from | What you supply | Band ceiling | Why |
|---|---|---|---|
| An imported TID-CMM assessment | A tid-cmm/export JSON document, or the detection score by hand | None — assessable to L5 | Detection maturity has been assessed against its own model, with its own constraints and evidence rule already applied. |
| The built-in prerequisite check | Eight questions — five on detection, three on threat modelling — of which at least six must be answered | L4 | A self-assessment is sufficient to assess response up to L4. Claiming adaptive response maturity on an unaudited detection claim is not credible. |
| Nothing | No import, and fewer than six prerequisite questions answered | L3 | The response score rests on an unaudited claim about visibility. It can still be useful; it cannot be presented as validated. |
Only the detection questions feed R4. The threat modelling questions produce a separate figure that indicates how much confidence your lattice scoping carries — an attack-path declaration made without any modelling behind it should not earn the same scope quality as an imported modelled one.
Answer the prerequisite check honestly and it will usually be less flattering than the TID-CMM import you do not have. That is the intended direction. A proxy that scores you higher than the audited assessment would is not a proxy, it is a wish.
The eight prerequisite questions
Each is scored on the same 0–5 scale as everything else in the model. The anchors below are the 3 anchor for each question — the honest middle, and the one most organisations should be scoring against.
| ID | Question | What a 3 looks like |
|---|---|---|
| PD-1 | Telemetry coverage of crown jewels | We know our log sources and they cover our crown-jewel systems. |
| PD-2 | Where detection content comes from | A mix of vendor content and rules we wrote for our own environment. |
| PD-3 | Detection testing | We test significant detections when we build them. |
| PD-4 | Alert quality reaching a responder | Alerts arrive enriched with asset and identity context. |
| PD-5 | Identity and cloud detection coverage | We have working detection for identity and cloud control-plane abuse. |
| ID | Question | What a 3 looks like |
|---|---|---|
| PM-1 | Crown jewel definition | An agreed, documented crown-jewel list owned by the business. |
| PM-2 | Attack path modelling | We have modelled the main attack paths to our critical systems. |
| PM-3 | Adversary prioritisation | We have named the threat actors and scenarios most likely to target us. |
The detection figure R4 uses is the mean of the detection questions you answered; the threat modelling figure is the mean of the modelling questions. Six answers of the eight are the minimum for the check to count at all — below that the tool falls back to supplying nothing, and the L3 ceiling applies.
None of this is an argument against running TID-CMM. An organisation that runs both gets a detection figure that has survived its own evidence rule, and the L5 ceiling lifts. The point is narrower: the absence of a detection assessment should cost you the top of the scale, not the assessment.
Before you start
Nothing here is mandatory — the tool will run on declarations alone — but every item you bring moves the result from opinion towards evidence, and the constraints reward that.
Your TID-CMM export, if you have one
A tid-cmm/export JSON file populates step 1 in one action: crown jewels, modelled attack paths, priority actors and their breakout times, and the detection score.
It also supplies D for constraint R4, and it is the only source that leaves the band uncapped. Without an import, answer the prerequisite check instead: the report is flagged self-assessed and the band is capped at L4. Supply neither and the cap is L3.
Your crown jewel list
Crown jewels mapped to asset classes, with RTO and RPO where you have them. This is what assigns each in-scope lattice cell a criticality tier — T1 where the cell is or hosts a crown jewel, T2 where it lies on a modelled path to one, T3 where it is peripheral.
Tiering drives the weighted score, the blind-cell gate and the roadmap. Guessing here degrades all three.
Priority actor breakout times
Breakout time is the interval from initial foothold to first lateral movement, taken from threat intelligence or the UTIOM metrics calculator. Where an actor-specific figure does not exist, an industry baseline is used instead.
A baseline figure must be flagged estimated, and an L4 claim is never allowed on estimated tempo.
Timing data from your case system
You need three durations: MTTD, MTTDecide and MTTC. MTTD usually arrives with the TID-CMM import or from the SIEM. MTTC is normally recoverable from containment action timestamps.
MTTDecide is the one nobody has. See below.
MTTDecide is usually not instrumented — assume that, and plan for it
MTTDecide is the mean time from validated alert to authorised containment decision, and it is the measurement TIR-CMM contributes. Every maturity model in circulation folds decision latency into MTTR, where it disappears. Organisations that separate it routinely discover it exceeds MTTC by an order of magnitude.
The obstacle is mechanical rather than conceptual. Many case systems cannot distinguish “alert validated” from “containment authorised” without a process change — there is no timestamp between the two, because no one was ever asked to record one. The metric is only as good as the timestamp discipline behind it.
Do not let that stop the assessment. Estimate MTTDecide from a sample of recent incidents — walk the case timeline, the chat log and the approval trail for each, and take the interval from the moment the alert was confirmed real to the moment someone with authority said act. Then flag the figure as estimated in the report, and raise the instrumentation change as a finding: an explicit authorisation timestamp in the case system is a small process change that turns the model’s most diagnostic metric from an estimate into a measurement. Sub-capability RA-2 scores precisely this, and AO-6 scores whether case data is complete enough to compute MTTD, MTTDecide and MTTC without manual reconstruction.
An estimated MTTDecide still produces a usable Containment Margin and Tempo Ratio. A missing one does not: supply no tempo at all and constraint R5 caps the band at L3 and marks the report tempo unverified.
The ten steps
Each step takes something you supply and produces something the model needs. Nothing is asked for twice, and no step can be completed with information the previous steps did not establish.
This is the Baseline walkthrough. The badge on each step says which tier first asks for it: Pulse steps are in all three tiers, Baseline steps are in Baseline and Assurance, and the three Assurance steps are listed after the tenth.
- Pulse Import or declare context.
You choose: standalone — the default — or an imported TID-CMM assessment.
You supply: in standalone mode, the eight prerequisite questions, then crown jewels, attack paths and priority actors by hand. On import, a TID-CMM JSON export fills all of that in one action.
It produces: crown jewels, attack paths, priority actors, the detection score that feeds constraint R4, and the band ceiling that source justifies — none on import, L4 on the prerequisite check, L3 on nothing. - Baseline What you run.
You supply: asset class selection from A1–A8.
It produces: the active columns of the lattice. Classes absent from the environment are excluded entirely rather than scored as gaps. - Baseline What you must keep running.
You supply: crown jewel to asset class mapping, with RTO and RPO.
It produces: cell criticality tiers T1–T3. - Baseline Who targets you and how fast.
You supply: priority actors and their breakout times.
It produces: B(a), the denominator of every tempo calculation. - Baseline Lattice scoping.
You supply: confirmation of the auto-derived N/A cells.
It produces: the in-scope lattice, typically 24–38 cells of a possible 64. - Baseline Containment options.
You supply: per asset class, the containment primitives and graded tiers actually available.
It produces: input to the CE domain, and the raw material for judging cell status. - Baseline Authority map.
You supply: who may authorise what, when, and with what latency.
It produces: input to the RA domain. This is the step that surprises people. - Baseline Lattice status.
You supply: a Response Readiness Status of 0–3 for each in-scope cell, with evidence and a date.
It produces: the Validated Response Score, the engineered and proven rates, and the list of blind cells. - Baseline Capability scoring.
You supply: 58 sub-capabilities scored 0–5, each with evidence.
It produces: the eight domain scores. - Pulse Tempo and results.
You supply: MTTD, MTTDecide and MTTC.
It produces: the full report and the ranked roadmap.
The three Assurance steps
Assurance runs the same ten and inserts three more, each at the point where its answers are needed. Full detail for all three is on the model page.
- Assurance Telemetry attributes —
after the authority map, before lattice status.
You supply: eight attributes scored 0–5 against each in-scope asset class — collection, completeness, timeliness, integrity and normalisation, retention, queryability, health monitoring, access protection.
It produces: a telemetry mean per asset class and a named list of weak attributes, which is what tells you whether a lattice cell is limited by containment or by evidence. - Assurance Scenario validation —
after capability scoring, before tempo.
You supply: for each in-scope scenario, ten lifecycle stages scored 0–5, the date it was last exercised, its target recovery time and the recovery time actually observed.
It produces: pass or fail per scenario, plus a run-but-unmeasured count. A scenario passes only when every stage is above 1, it was exercised, the exercise was timed, and the observed recovery met its target. - Assurance Governance — after tempo,
before results.
You supply: yes or no against the eight assurance conditions, worked through with the ten calibration questions in hand.
It produces: the governance completeness figure and the R7 governance ceiling. Miss evidence-led scoring, independent calibration or separation of assessor from approver, and the result is a self-assessment capped at L4.
An exercise with no recorded times cannot lift a readiness claim. A scenario that was run but not timed counts as neither a pass nor a failure — it is reported as run-but-unmeasured. Book the stopwatch before you book the room: observed timings are what make a scenario proof rather than paperwork.
Why the authority map comes before the lattice
Step 7 is placed before step 8 deliberately, and the ordering does real work.
Ask an assessor to score containment readiness cold, and they score the tooling. The EDR can isolate a host, so the endpoint column looks engineered. The identity platform can revoke a session, so Tier-0 looks covered. What that answer measures is whether the button exists, not whether anyone is permitted to press it — and the model’s entire argument is that the button was never the constraint.
Mapping authority first changes the question being answered. By the time an assessor reaches step 8 they have just written down, action by action, which containment steps need a named approver, who that approver is, whether a deputy exists, and what the response time is out of hours. They have just discovered how many actions require an approval they cannot obtain at 03:00 on a Sunday. Cells that would have been scored 2 — engineered: documented, tooled, owned, authority defined in advance — are scored 1 instead, because the assessor now knows the action is gated behind an approval path with no defined SLA, which is the definition of status 1.
The result is a lattice scored against the organisation’s real capacity to act rather than against its product inventory. Reversing the two steps does not merely reorder the questionnaire; it systematically inflates the Validated Response Score, and it hides the finding — decision latency — that the model exists to surface.
Automation you are not permitted to fire is a demonstration, not a capability. Constraint R2 makes the same point arithmetically, capping AO and CE at RA + 1. Step ordering makes it before the arithmetic ever runs.
What the assessment gives back
A band and a decimal are not an answer to anything. Step 10 therefore produces four things beyond the score: three readiness figures leadership can act on, a plain-language account of the current state, a ranked set of findings, and a roadmap bucketed by how much effort each item actually takes.
Three readiness lenses
The eight domains describe how the capability is built. Leadership asks a different question — are we ready to respond, ready to recover, and able to keep running. The lenses re-cut the same 58 sub-capabilities against those three outcomes. They are lenses, not partitions: a sub-capability can serve more than one outcome, and eight of them do.
| Lens | The question it answers | What feeds it | Subs |
|---|---|---|---|
| Readiness to respond | If it started now, could we act on it in time? | Authority to act, playbooks that exist and run, containment options with known blast radius, and the automation and investigation depth to use them. | 27 |
| Readiness to recover | Could we get the business back, clean, and prove it? | Eradication that actually removes the adversary, restoration with integrity verification, backups that survive a destructive attack, and the legal and communications machinery that runs alongside. | 14 |
| Operational resilience | Could we keep running through it, and be harder to hit next time? | Degraded-mode operation, a response function that survives losing its own tooling or people, and a feedback loop that converts every incident into a structural improvement. | 15 |
Each lens reports two figures. The claimed figure is the weighted mean of its sub-capabilities as scored. The proven figure applies the same rehearsal ceiling as everything else in the model — validation maturity plus one — because unrehearsed readiness is assumed readiness, and a lens is not exempt from that.
This produces an effect worth knowing about before you read the output. When the rehearsal ceiling binds, it flattens all three proven figures to the same number. Three identical readiness scores are not a finding about your response, recovery and resilience being equally good; they are a finding about your exercising. In that state it is the claimed figures that discriminate between the three, and the gap between claimed and proven is itself the headline — it is the size of the part of your readiness that has never been executed.
The current-state report
Written prose, generated from your own numbers, in language that survives being pasted into a board paper without a glossary. It states the band and the score, and the difference between the self-assessed and adjusted figures — which is the part of the capability that is currently assumed rather than demonstrated. Where a band cap is doing the work rather than the arithmetic, it says so plainly, because that changes what improvement will and will not move.
It then reports the strongest and weakest lens and whether either is capped, the engineered and proven rates across the lattice with any crown-jewel cell that has no response option named individually, and the tempo position: containment time against breakout time, and where the minutes actually go. Supply no tempo at all and it says that too — the most decisive question the model asks is then unanswered, and the report should not pretend otherwise.
What is actually holding you back
Alongside the prose, a ranked set of findings, each with a severity and a what to do. They are derived mechanically, not written for you, and they come from four places:
- Binding constraints — R1, R2 and R4, translated out of model language, with the number of domains each pulled down and the points it cost.
- Band caps — R5 where tempo is losing or unmeasured, R6 where a crown-jewel stage has no response option at all.
- The weakest domains — the lowest three, reported only where they score below 3, with the domain’s own core question as the thing to fix.
- The coverage gap — raised where the proven rate is low and engineered coverage runs well ahead of it, because that difference is the honest measure of how much of the capability is an assumption.
The finding that most often surprises people looks like this:
Your tooling has outrun your authority to use it
Automation and containment are capped by your decision-rights score. This is the failure that looks like a technology problem and is not one.
No purchase will move this. Written, signed pre-authorisation from named business owners will.
The improvement blueprint
Every one of the 58 sub-capabilities carries a concrete action — imperative, startable on Monday — together with an effort tier, a type (authority, process, engineering, exercise, measurement, tooling or governance), the role that must actually do it, and what you hold once it is done. The roadmap ranking stays mechanical; the blueprint buckets it by effort, so the output separates what you could do before Friday from what needs a budget line.
| Horizon | Window | What it means | Items |
|---|---|---|---|
| Quick fixes | 0–30 days | One person, no budget, no procurement. Mostly writing things down, getting decisions made, and measuring what you already do. | 20 |
| Ninety-day moves | 1–3 months | Needs coordination, a scheduled exercise, or a contained piece of engineering. Achievable inside a quarter with existing people. | 23 |
| Structural work | 6–12 months | Needs budget, procurement, hiring, architecture change or a standing programme. Plan it now; it will not land this quarter. | 15 |
Gaps found on the lattice are bucketed too. A blind crown-jewel cell is a quick fix, because naming an owner and writing an interim manual procedure is a days-not-quarters job and it is what lifts the R6 band cap — even a bad documented option beats none. Every other lattice gap is a ninety-day move: build or document the containment action, and name who may authorise it.
Authority and measurement work is overwhelmingly quick-fix effort. Of the ten sub-capabilities whose action is an authority or a measurement one, seven are quick fixes and none are structural. Which is the uncomfortable finding hiding inside the roadmap: the highest-leverage items in this model — the ones that release R2, that make MTTDecide visible, that put a named person on the 03:00 decision — usually cost nothing but a decision.
Running it continuously
The three tiers describe how deep a single assessment goes. They say nothing about how often you run one, and that is a separate decision with a clear answer.
Continuous is the cadence the model is built for. Maintain lattice status live, automate the validation expiry, and reassess quarterly — deepening a tier at a time rather than re-running the same tier.
| Cadence | What it produces | What it misses |
|---|---|---|
| One-off | A position, and a roadmap that is accurate on the day it is written. | Every expiry. Cells silently revert from proven to engineered and nobody is told. |
| Annual | A trend line, provided scope and method are comparable year to year. | The twelve-month expiry lands between assessments, so a year’s validation debt arrives as a surprise. |
| Continuous | A standing validation-debt figure, deepening tier by tier as evidence accumulates. | Nothing structural — but it needs an owner, which is what RG-7 scores. |
The reason cadence matters this much is specific to response. Response maturity is perishable in a way detection maturity is not: a cell at status 3 reverts to 2 after twelve months, or sooner on an EDR swap, a SOAR migration, an identity platform change, a change to the on-call or escalation model, a reorganisation of the response function or the departure of the named owner, a material architecture change, or a change of the managed provider delivering the action. Maintained live, that expiry is a standing validation-debt figure. Run once a year, it is a surprise.
Scoring honestly
Constraint R3 — the evidence cap is the only constraint an assessor can trigger directly, and it is the one that decides whether the assessment is worth reading. Every score you enter is accompanied by an evidence level of 0 to 3, and that level caps the score. Unsupported claims must not outrank evidenced assessments.
The cap applies mechanically, before any domain score is computed, and it cannot be argued with. In the model’s numerical validation it bound six sub-capabilities — preparation items and, tellingly, pre-authorised containment actions — all held at design-or-policy evidence and therefore capped at 2.
Choosing the level
Ask what you would hand to somebody who did not believe you.
- Level 0 — assertion only. Nothing, or a verbal claim. Caps the score at 1. This is the level a blank answer is treated as, so leaving it empty is not a way of avoiding the question.
- Level 1 — design or policy. A document, a policy, a screenshot, a partial implementation. Caps at 2. Something exists on paper; whether it works is unproven.
- Level 2 — implemented and tested. Implementation evidence plus a test result, a ticket or a system record. Caps at 4. This is where a named, dated artifact lands: an incident ticket ID, an exercise report with a date, a signed authorisation matrix mapped to lattice cells, a timestamped containment log.
- Level 3 — repeatably validated. Recent, repeatable, independently reviewable proof from the live environment. Caps at 5. Not one successful test — a standing practice somebody outside your team could re-run.
Elsewhere in the model the same standard appears as the evidence column against every sub-capability: an unannounced call-out test record, a metric series with timestamps drawn from the case system, a rollback test record, a restore test record, a live-fire schedule with results per cell, linked incident-to-detection-change records. Each is nameable and each is dateable.
What does not count
“We do that” is level 0. A policy asserting that something must happen is level 1, not level 2 — a policy is a statement of intent, and the cap exists precisely to stop intent scoring as delivery. A slide describing a capability, a team’s shared confidence that the procedure would work, and a document with no version and no date are all level 0.
A capped score is not a failure. It is the correct score for a capability that is real, consistently practised and undocumented in evidence terms. The model reserves the top of the scale for capability that has left a trace.
The cap is a ceiling, not a conversion, and it is the finding. A high evidence level permits a high score; it does not create one. Level 3 evidence against a capability you score 2 leaves the score at 2. And raising a level to lift a cap does not improve the capability — it deletes the most useful thing the assessment found. If you are tempted, write the gap down instead.
The one place none of this applies is Pulse, which does not ask for evidence levels at all. Twenty minutes is not enough time to review artifacts honestly, so the cap is switched off and the reduced rigour is carried by the L3 band ceiling under R7 instead.
The same discipline applies to lattice cells. Status 3 — proven — requires execution against a real incident or a live-fire exercise inside the recency window, meeting the stage time objective, at the designed blast radius. All three conditions, with a date. Status 2 is engineered and unproven, and the gap between the two rates is the headline finding of the whole assessment.
Privacy
The assessment tool is entirely client-side. There are no server calls, no analytics and no transmission of anything you enter. Your crown jewels, your attack paths, your authority map, your incident timings and your scores stay in your browser.
This is inherited from TID-CMM and it is non-negotiable — it is the reason practitioners will run the model on real data rather than on a sanitised version of their estate. An assessment run on sanitised inputs produces a sanitised result, which is worse than no result at all, because it looks like a finding.
You export when you choose to, as a JSON file that lands in your downloads folder and goes wherever you send it — to UTIOM’s roadmap tool, to your own records, to nowhere.