Data Centre Criticality Matrix
A criticality matrix converts opinions about which equipment matters into a score that can be applied consistently across a site. Without one, maintenance frequency, spares holding and response priority drift towards whichever asset was most recently troublesome.
Scope note
This article provides general operational guidance. Apply site-specific engineering review, risk controls, manufacturer requirements and applicable regulations before use.
Overview
This article is about the artefact: the dimensions you score, the bands you define, how three example assets score against them, and what each band should change in practice. The wider method — how criticality is governed, challenged and used in asset decisions — sits on the Data Centre Asset Criticality page, which parents this one.
This article provides general operational guidance. Apply site-specific engineering review, risk controls, manufacturer requirements and applicable regulations before use.
What a criticality matrix does
The matrix has one job: to produce a repeatable ranking of assets based on the consequences and exposure associated with their loss, so that finite maintenance, spares and engineering attention are allocated against consequence rather than familiarity.
It is a decision input, not a decision. A score does not authorise or prevent work; it sets the default treatment for an asset and identifies the cases where a deliberate exception should be recorded. Treat the matrix as a live governed document with an owner, a revision and a review interval — the same controls you would apply to the asset register itself.
- Rank by consequence of loss to the service, not by equipment type or capital value.
- Apply one scale across the whole site so results can be compared.
- Keep the scoring rules written down; an unstated rule is applied differently by each assessor.
- Record the reasoning for each score, not only the number — the reasoning is what gets challenged at review.
Choosing the scoring dimensions
Most workable matrices use a small number of dimensions. Adding more feels rigorous but usually produces the same ranking with more effort and more disagreement. The set below is a common starting point; adapt the dimensions and weights to your own service commitments and engineering review.
Two of these dimensions are often confused. Service consequence describes what is lost if the asset fails. Redundancy exposure describes how likely that loss is to reach the service — an asset in a resilient configuration can have severe consequence but low exposure. Keeping them separate is what stops every item in a critical system scoring identically.
The scoring scales below are an illustrative starting structure, not an industry standard or prescribed data-centre criticality methodology. Site-specific weights, thresholds and treatment should be established through engineering and operational review.
| Dimension | What it measures | Typical scale | Common evidence source |
|---|---|---|---|
| Service consequence | Effect on IT load, cooling or life safety if the asset fails | 1 (none) – 5 (immediate load loss) | Single-line diagrams, system design intent |
| Redundancy exposure | Whether a failure is absorbed by the configuration | 1 (2N absorbed) – 5 (single path, no alternative) | Topology, current configuration, concurrent works |
| Detectability | Whether degradation is visible before failure | 1 (monitored and trended) – 5 (hidden until it fails) | BMS/EPMS points, monitoring coverage, inspection regime |
| Recovery time | Time to restore function after failure | 1 (minutes, switchable) – 5 (long lead item) | Spares holding, supplier lead time, prior repairs |
| Safety and compliance | Life-safety or statutory function involved | 1 (none) – 5 (life-safety or statutory duty) | Fire and life-safety design, statutory inspection register |
Illustrative total-score bands
With five dimensions each scored 1–5, totals run from 5 to 25. The thresholds below show one way of collapsing those totals into the four bands used in the rest of this article.
These thresholds are illustrative. If the site uses weighted dimensions, different scoring scales or different treatment thresholds, document those rules in the approved matrix rather than changing scores retrospectively to achieve a desired band.
| Total score | Band |
|---|---|
| 17–25 | C1 — Critical |
| 13–16 | C2 — Essential |
| 9–12 | C3 — Supporting |
| 5–8 | C4 — Non-critical |
Score reference · 5–25
Illustrative criticality matrix
C4
Non-critical
5–8
C3
Supporting
9–12
C2
Essential
13–16
C1
Critical
17–25
Illustrative colour treatment
Use the colour treatment as a visual aid when reviewing the matrix. C1 Critical assets fall in the red range, C2 Essential assets in the orange range, and C3 Supporting/C4 Non-critical assets in the green range. The numerical score and approved C1–C4 treatment rules remain the controlling classification.
Colour is a visual aid; the numerical score and approved C1–C4 treatment rules remain the controlling classification. Green does not mean zero risk.
Defining consequence bands
Scores are only useful once they collapse into a small number of bands with defined treatment. Three or four bands is usually enough; beyond that the distinction between adjacent bands stops being defensible.
Define each band by what it commits the organisation to do, not by adjective. 'High' means nothing operationally; 'quarterly inspection, spares held on site, response within the shift' means something. Write the band definition before scoring anything, or the bands will be quietly adjusted to suit the results. The numerical thresholds shown above are illustrative: each organisation should establish and approve its own thresholds and treatment rules.
| Band | Description | Typical treatment | Review frequency |
|---|---|---|---|
| C1 — Critical | Failure causes immediate loss or unprotected exposure of IT load, cooling or a life-safety function | Highest maintenance frequency, condition monitoring where practicable, spares held, MOP-controlled intrusive work | Annually and after any configuration change |
| C2 — Essential | Failure removes resilience or degrades service but is absorbed by the configuration | Planned maintenance at defined intervals, defined spares strategy, defect escalation path | Annually |
| C3 — Supporting | Failure has limited operational effect and can be scheduled for repair | Routine maintenance, spares by exception, standard defect handling | On change or at longer interval |
| C4 — Non-critical | No operational or safety effect within the service boundary | Minimum compliant maintenance | On change |
Redundancy and exposure factors
Design-state criticality is not the same as operational-state exposure. An asset may be adequately redundant under normal configuration but temporarily become a single point of exposure during maintenance, isolation or concurrent works.
Redundancy reduces exposure to a failure, but it does not reduce the consequence of losing the function, and it is not permanent. An N+1 system running with one unit on maintenance is, for that period, a single-path system. A matrix that scores against the design topology alone will understate risk during exactly the periods when work is being performed.
Handle this by scoring against the design configuration and recording the degraded-state effect alongside it, rather than rescoring the asset every time plant is taken out of service. The degraded-state note is what maintenance planning and window approval should read when concurrent works are being assessed.
- Score exposure against the design configuration, and record separately what the band becomes when redundancy is unavailable.
- Treat shared upstream infrastructure as higher exposure than the individual units it feeds.
- Check whether claimed redundancy is proven — an untested alternative path is a design intent, not a demonstrated capability.
- Where a single component serves multiple redundant paths, score it on the paths it can remove together.
Worked scoring of three example assets
The following assets are illustrative and use the example dimensions above with equal weighting. They demonstrate how the scoring structure can separate assets that are often treated as equally important; they are not a recommendation for any specific site or a prescribed industry standard.
The UPS module and the CRAH unit both sit in resilient systems, but they separate on consequence, recovery time and safety exposure. The lighting circuit scores low on service consequence but does not fall to the lowest band because of its safety-related function — which is the outcome the safety dimension exists to produce.
- UPS module: consequence is high because the protected load depends on the function; recovery time carries the score because module replacement is not a same-shift activity at most sites.
- CRAH unit: the configuration absorbs a single loss, but exposure rises while other units are unavailable and repair typically runs beyond a single shift, placing it at the lower end of C2.
- Lighting circuit: low service consequence, but the safety dimension prevents it dropping to the bottom band where emergency or plant-room lighting is involved.
| Asset | Service consequence | Redundancy exposure | Detectability | Recovery time | Safety / compliance | Total | Band |
|---|---|---|---|---|---|---|---|
| UPS module in an N+1 system | 5 | 3 | 2 | 4 | 3 | 17 | C1 |
| CRAH unit in a hall with N+2 cooling | 4 | 3 | 2 | 3 | 1 | 13 | C2 |
| Lighting circuit serving a plant room | 1 | 2 | 3 | 1 | 3 | 10 | C3 |
Turning scores into maintenance and spares decisions
The matrix earns its place only when the band changes something. If C1 and C3 assets receive the same maintenance regime, the same spares treatment and the same defect priority, the scoring exercise has produced a document rather than a control.
Connect each band to a default in the maintenance plan, the spares policy and the defect-handling process, and require an exception note where an asset is treated differently from its band. Exceptions are legitimate — access constraints, supplier arrangements, planned replacement — but they should be visible and owned rather than accidental.
| Band | Maintenance regime | Spares holding | Defect priority | Review frequency |
|---|---|---|---|---|
| C1 | Highest planned frequency plus condition monitoring where practicable; intrusive work under a MOP | Critical spares held or contractually guaranteed | Immediate assessment; escalation if resilience is affected | Annual and on change |
| C2 | Planned maintenance at defined intervals | Spares strategy defined; long-lead items identified | Scheduled within a defined period | Annual |
| C3 | Routine maintenance | Sourced on demand unless lead time is long | Planned repair | Longer interval or on change |
| C4 | Minimum compliant maintenance | None held | Batched with planned works | On change |
Keeping the matrix current
Criticality changes when the facility changes. Load growth, a new client requirement, a topology change, a temporary configuration that becomes permanent, or the removal of a redundant path all move assets between bands — usually without anyone revisiting the matrix.
Tie the review to events as well as to the calendar, and hold the matrix against the asset register so that new, modified and decommissioned assets are reflected. A matrix that no longer matches the register is the most common failure mode, and it is a document-control problem rather than an engineering one.
- Re-assess after a configuration or topology change, a load change, or a change to service commitments.
- Re-assess after a failure that had a larger effect than the band predicted — that is evidence the scoring was wrong.
- Reconcile against the asset register at each review; unmatched records in either direction are defects.
- Record the owner, the revision and the date of each assessment on the matrix itself.
Key takeaways
- 01Score consequence and exposure separately, or every asset in a critical system scores the same.
- 02Define what each band commits you to before scoring anything.
- 03Redundancy reduces exposure, not consequence — and it disappears during maintenance.
- 04Detectability and recovery time usually separate assets that look equally critical.
- 05A band must change maintenance, spares and defect priority, or the matrix is decoration.
- 06Re-assess on change and after any failure worse than the band predicted.
- 07Reconcile the matrix against the asset register at every review.
Frequently asked questions
- What is asset criticality?
- Asset criticality is a ranking of equipment by the consequence of its loss to the service the facility provides — not by capital value, size or age. In a data centre that usually means the effect on IT load, cooling, and life-safety or statutory functions, adjusted for whether the configuration would absorb the failure.
- How many criticality bands should I use?
- Three or four is usually enough. Each band must be distinguishable by what it commits you to do — maintenance frequency, spares holding, defect priority, review interval. If two adjacent bands lead to the same treatment, merge them.
- How often should criticality be reassessed?
- Set a periodic interval in site governance, and in addition re-assess on trigger events: configuration or topology changes, load or client-requirement changes, plant replacement, and any failure whose effect exceeded what the band predicted.
- Does redundancy reduce criticality?
- It reduces exposure to a single failure; it does not reduce the consequence of losing the function, and it is temporarily absent whenever a unit is under maintenance. Score the design configuration and record separately what the band becomes in the degraded state.
Free resource
Data Centre Asset Criticality Matrix
Structure a consistent first-pass view of service consequence and asset importance.
Get the checklistProfessional toolkit
Data Centre Asset Management Toolkit
Practical tools for asset criticality, condition, risk, lifecycle planning and register governance.
US$59
Coming soon — not available for purchase
View product