Critical Ops HubOperations Resource
MCOPS-PRO-022Rev 1.0
MOP / SOP / EOP

Data Centre MOP Review Checklist

A method of procedure (MOP) in a data centre is a controlled document describing a planned intervention on live infrastructure. Reviewing one is not proofreading. The reviewer is deciding whether the written sequence will keep the facility inside its acceptable operating envelope when a technician follows it literally at three in the morning.

By Critical Ops HubPublished 23 March 2026Updated 23 May 20267 min read

Scope note

This article provides general operational guidance. Apply site-specific engineering review, risk controls, manufacturer requirements and applicable regulations before use.

Overview

This checklist works through the review in the order risk accumulates: scope, affected systems, prerequisites, risk and operational impact, isolation and switching, controls and dependencies, communications, rollback, hold points, testing and verification, authorisation, contingency, post-work verification and documentation.

It includes the reviewer's walkthrough: reading the procedure as the executing technician rather than as its author. A technically correct MOP is not necessarily an operationally complete MOP. It may describe the engineering action accurately while omitting the plant state, decision rights, communications, recovery time or evidence needed to control that action in a live facility.

Technical correctness is only one approval test

Technical review asks whether the proposed actions can achieve the intended engineering result. Operational review asks whether those actions can be authorised, controlled, observed, interrupted, recovered and closed out under the actual site conditions. Approval needs both answers.

For example, a switching sequence may identify the correct devices in the correct order but still be incomplete if it assumes the alternate path is healthy, leaves the control-room alarms unexplained, has no hold before loss of redundancy, or gives no safe state for an overrun. The reviewer should therefore test the MOP as a control system, not only as a sequence of technical instructions.

1. Scope

Scope failures are the most common reason a MOP is rejected. The document must state precisely what is being worked on, what is explicitly excluded, and where the boundary of the activity sits.

  • Single, unambiguous statement of the work to be performed.
  • Named assets with their register references, not generic descriptions.
  • Explicit exclusions — what this MOP does not cover.
  • Physical and system boundaries of the intervention.
  • Reject if the scope could be read two ways by two competent technicians.

2. Affected systems and prerequisites

The affected-systems list must extend past the asset being worked on to everything downstream of it and everything that shares a dependency with it — control systems, monitoring, fire detection interlocks and cooling that depends on the same power path.

Prerequisites are the conditions that must be true before step one. They belong in a list the executing team can verify on the day, each with a pass criterion.

Prerequisite review
Prerequisite typeWhat the reviewer checksReject if
Plant stateRequired redundancy is available and provenState is assumed rather than verified on the day
DocumentationDrawings and schematics current and referenced by revisionReferenced by title only
CompetencyNamed roles with required competenciesRole stated as 'technician'
PPE and safety equipmentTask-specific PPE and safety equipment listed, available and fit for purposePPE or task-specific safety equipment is unavailable, damaged, unsuitable or outside a required inspection or calibration period where applicable
EquipmentTest equipment listed with calibration requirementCalibration not addressed
EnvironmentConcurrent works and freeze periods checkedNo concurrency check

3. Risk assessment and operational impact

The review confirms that the risk assessment addresses the operational consequence of the work, not only the personnel safety consequence. Both matter; only one of them is usually documented well.

Operational impact should be stated in terms the operations team can act on: which systems lose redundancy, for how long, what the facility's tolerance is during that period, and what condition would make the risk unacceptable.

  • Worst credible operational outcome stated, not just the expected outcome.
  • Duration of any reduced-resilience state quantified.
  • Concurrent-risk conditions identified (weather, other works, elevated load).
  • Stop conditions written as objective criteria.

4. Isolation, switching sequence and controls

Every switching step should be individually numbered, individually verifiable and individually reversible. The reviewer's test is simple: can a competent technician perform this step, confirm the expected result, and know what to do if that result does not appear?

Three rewrite examples make the standard concrete. 'Isolate the UPS' is rejected — no device, no method, no confirmation. 'Open UPS-2 input isolator IS-2A and confirm open by local indication' is accepted. 'Switch cooling to manual' is rejected — no setpoint, no target state. 'Place CRAC-04 into manual at setpoint 24 °C and confirm the unit reports manual mode at the BMS' is accepted. 'Check everything is normal' is rejected — not a test. 'Confirm the downstream board reads within its normal voltage band as recorded in the pre-work readings' is accepted.

Step anatomy the reviewer should find on every switching step
ElementRequirement
ActionOne action, one device, named by its unique reference
MethodHow the action is performed and by whom
Expected resultThe observable indication that confirms it worked
ToleranceThe acceptable range where the result is a measurement
If not obtainedExplicit instruction — stop, hold, revert, escalate

5. Dependencies and communications

Dependencies are the steps that cannot proceed without something outside the executing team: a client notification, a vendor attendance, a load transfer completed elsewhere, a weather condition. Each should be named with its owner and its confirmation point.

The communications plan needs to state who is informed before, during and after, through which channel, and who makes the call if communications fail.

  • Named owner for every external dependency.
  • Pre-start, in-progress and completion notifications defined.
  • Single named point of contact during execution.
  • Defined behaviour if the control room cannot be reached.

6. Rollback, recovery and hold points

A rollback plan is only credible if it is still achievable at the point of no return. The reviewer should walk the sequence and identify the step after which the original state cannot be restored within the accepted window — then confirm that a hold point sits immediately before it.

Hold points must name who acknowledges them. A hold point without an owner is a pause, not a control.

  • Rollback written as a sequence, not as 'reverse the above'.
  • Time required for rollback stated and compared against the accepted window.
  • Point of no return identified explicitly.
  • Hold point before every irreversible step, with a named acknowledger.
  • Recovery actions for the credible failure modes, not only for success.

7. Testing, verification and authorisation

Testing confirms the work achieved its purpose. Verification confirms the facility is back in its intended state. They are separate checks and reviewers frequently find only the first.

Authorisation should name the approving roles and record that each of them reviewed the document at its current revision. A signature against a superseded revision invalidates the review.

Verification review
CheckReviewer confirms
Functional testTest proves the intended outcome, with a pass criterion
Return to normalEvery altered setting, mode, inhibit and isolation has a removal step
Redundancy proofRestored redundancy is demonstrated, not assumed
Baseline comparisonPost-work readings compared against recorded pre-work readings
AuthorisationCorrect roles, current revision, dated

The 26-point MOP review check

Use the following as the final review register. Record Pass, Fail or Comment against every line. A failure is not cured by a general approval note; the MOP should be revised, reissued and reviewed at the new revision.

Check 1 is a readiness gate. Before working through the remaining checks, confirm that all task-specific PPE and safety equipment required by the MOP is available, fit for purpose, undamaged, within any required inspection or calibration period where applicable, and ready for use. If it is unavailable, unsuitable for the task, damaged, not fit for purpose, outside a required inspection or calibration period as required by the applicable site or task requirement, or otherwise not ready for safe use, the MOP is not ready to proceed and the review should be recorded as a fail against Check 1 pending resolution.

MOP approval register
#Review checkPass / Fail / Comment
01Task-specific PPE and safety equipment required by the MOP is available, fit for purpose, undamaged and within required inspection or calibration periods where applicable
02Scope, exclusions and work boundary are unambiguous
03Assets use current unique references
04Affected upstream, downstream and shared systems are identified
05Required initial plant state is explicit
06Current drawings and references carry revisions
07Competent roles and supervision are named
08Tools, spares and calibrated test equipment are confirmed
09Concurrent work and degraded states are checked
10Operational impact and reduced resilience are stated
11Worst credible outcome is addressed
12Objective stop conditions are defined
13Isolation boundary and proving method are clear
14Each switching step names one action and one device
15Every critical action has an expected result
16Measured results have an acceptance range or criterion
17If-not-obtained instructions stop improvisation
18Dependencies have owners and confirmation points
19Communications cover pre-start, execution and completion
20Hold points precede irreversible or high-exposure steps
21Rollback is a timed, executable sequence
22Recovery covers credible failure modes
23Testing proves the work objective
24Verification proves restoration and redundancy
25Approvers, distribution and document revision are controlled
26Close-out records, defects and handover are assigned

Approval, revision and controlled distribution

The approval block should distinguish technical review from operational authority. The technical reviewer confirms the method and engineering logic; the operations authority accepts the site state, timing and residual operational exposure. Other approvals may be required by the site's governance, but a signature should never obscure what that person is accountable for.

The issue used at the workface must match the approved revision. Withdraw superseded copies, identify the controlled execution copy, brief every executing person on changes, and ensure the control room and work party refer to the same step numbering. If the method changes after approval, stop and send the revision back through the defined review route.

8. Contingency, post-work verification and documentation

Contingency covers what happens when the work cannot be completed: overrun beyond the accepted window, a failed component that cannot be replaced tonight, or a system that will not return to its normal state. Each needs a defined safe holding condition.

Documentation closes the loop — completed procedure, readings, defects raised, asset updates and the handover entry for the next shift.

  • Defined safe holding state if the work must be abandoned mid-sequence.
  • Overrun escalation route with a named decision maker.
  • Post-work verification signed by operations, not only by the executing team.
  • Defects and observations routed to the register on the day.
  • Completed procedure returned to document control at its issued revision.

The reviewer's walkthrough

Before signing, read the procedure once as the executing technician: standing in front of the plant, at night, with no access to the author. At every step ask what device is being touched, how the result is confirmed, and what happens if it does not appear.

Steps that survive that reading are executable. Steps that require the author's intent to interpret are not, however technically correct they may be.

Key takeaways

  • 01Reject any scope that two competent technicians could read differently.
  • 02Check affected systems past the asset itself — dependencies, controls and monitoring.
  • 03Every switching step needs an action, expected result, tolerance and an if-not-obtained instruction.
  • 04Identify the point of no return and put a named hold point in front of it.
  • 05Testing and verification are separate checks; confirm both exist.
  • 06Read the procedure as the technician before you sign it as the reviewer.

Frequently asked questions

Who approves a data centre MOP?
At minimum, a competent technical reviewer should confirm the method and an accountable operations authority should accept the operational state and residual exposure. Site governance may require additional approvals.
How long before the work should a MOP be reviewed?
Early enough to resolve comments, issue a controlled revision and brief the work party before the maintenance window. The correct lead time depends on complexity, risk and site change-control requirements.
What are the most common reasons a MOP is rejected?
Ambiguous scope, assumed plant state, unverifiable steps, missing hold points, weak rollback, undefined stop conditions, incomplete restoration tests and uncontrolled revision status.
Does every task need a MOP?
Not necessarily. The site’s work-control rules should determine when a task requires a MOP, SOP, work instruction or another control. Risk, complexity, system impact and repeatability should inform that decision.

Free resource

MOP Review Checklist

Review scope, prerequisites, step quality, hold points and recovery arrangements before approval.

Get the checklist

Professional toolkit

MOP / SOP / EOP Toolkit

Controlled templates and review tools for maintenance, standard operations and emergency response procedures.

US$59

Coming soon — not available for purchase

View product

Continue reading

Related articles

Back to the MOP / SOP / EOP pillar →