Data Centre MOP Review Checklist
A method of procedure (MOP) in a data centre is a controlled document describing a planned intervention on live infrastructure. Reviewing one is not proofreading. The reviewer is deciding whether the written sequence will keep the facility inside its acceptable operating envelope when a technician follows it literally at three in the morning.
Scope note
This article provides general operational guidance. Apply site-specific engineering review, risk controls, manufacturer requirements and applicable regulations before use.
Overview
This checklist works through the review in the order risk accumulates: scope, affected systems, prerequisites, risk and operational impact, isolation and switching, controls and dependencies, communications, rollback, hold points, testing and verification, authorisation, contingency, post-work verification and documentation.
It includes the reviewer's walkthrough: reading the procedure as the executing technician rather than as its author. A technically correct MOP is not necessarily an operationally complete MOP. It may describe the engineering action accurately while omitting the plant state, decision rights, communications, recovery time or evidence needed to control that action in a live facility.
Technical correctness is only one approval test
Technical review asks whether the proposed actions can achieve the intended engineering result. Operational review asks whether those actions can be authorised, controlled, observed, interrupted, recovered and closed out under the actual site conditions. Approval needs both answers.
For example, a switching sequence may identify the correct devices in the correct order but still be incomplete if it assumes the alternate path is healthy, leaves the control-room alarms unexplained, has no hold before loss of redundancy, or gives no safe state for an overrun. The reviewer should therefore test the MOP as a control system, not only as a sequence of technical instructions.
1. Scope
Scope failures are the most common reason a MOP is rejected. The document must state precisely what is being worked on, what is explicitly excluded, and where the boundary of the activity sits.
- Single, unambiguous statement of the work to be performed.
- Named assets with their register references, not generic descriptions.
- Explicit exclusions — what this MOP does not cover.
- Physical and system boundaries of the intervention.
- Reject if the scope could be read two ways by two competent technicians.
2. Affected systems and prerequisites
The affected-systems list must extend past the asset being worked on to everything downstream of it and everything that shares a dependency with it — control systems, monitoring, fire detection interlocks and cooling that depends on the same power path.
Prerequisites are the conditions that must be true before step one. They belong in a list the executing team can verify on the day, each with a pass criterion.
| Prerequisite type | What the reviewer checks | Reject if |
|---|---|---|
| Plant state | Required redundancy is available and proven | State is assumed rather than verified on the day |
| Documentation | Drawings and schematics current and referenced by revision | Referenced by title only |
| Competency | Named roles with required competencies | Role stated as 'technician' |
| PPE and safety equipment | Task-specific PPE and safety equipment listed, available and fit for purpose | PPE or task-specific safety equipment is unavailable, damaged, unsuitable or outside a required inspection or calibration period where applicable |
| Equipment | Test equipment listed with calibration requirement | Calibration not addressed |
| Environment | Concurrent works and freeze periods checked | No concurrency check |
3. Risk assessment and operational impact
The review confirms that the risk assessment addresses the operational consequence of the work, not only the personnel safety consequence. Both matter; only one of them is usually documented well.
Operational impact should be stated in terms the operations team can act on: which systems lose redundancy, for how long, what the facility's tolerance is during that period, and what condition would make the risk unacceptable.
- Worst credible operational outcome stated, not just the expected outcome.
- Duration of any reduced-resilience state quantified.
- Concurrent-risk conditions identified (weather, other works, elevated load).
- Stop conditions written as objective criteria.
4. Isolation, switching sequence and controls
Every switching step should be individually numbered, individually verifiable and individually reversible. The reviewer's test is simple: can a competent technician perform this step, confirm the expected result, and know what to do if that result does not appear?
Three rewrite examples make the standard concrete. 'Isolate the UPS' is rejected — no device, no method, no confirmation. 'Open UPS-2 input isolator IS-2A and confirm open by local indication' is accepted. 'Switch cooling to manual' is rejected — no setpoint, no target state. 'Place CRAC-04 into manual at setpoint 24 °C and confirm the unit reports manual mode at the BMS' is accepted. 'Check everything is normal' is rejected — not a test. 'Confirm the downstream board reads within its normal voltage band as recorded in the pre-work readings' is accepted.
| Element | Requirement |
|---|---|
| Action | One action, one device, named by its unique reference |
| Method | How the action is performed and by whom |
| Expected result | The observable indication that confirms it worked |
| Tolerance | The acceptable range where the result is a measurement |
| If not obtained | Explicit instruction — stop, hold, revert, escalate |
5. Dependencies and communications
Dependencies are the steps that cannot proceed without something outside the executing team: a client notification, a vendor attendance, a load transfer completed elsewhere, a weather condition. Each should be named with its owner and its confirmation point.
The communications plan needs to state who is informed before, during and after, through which channel, and who makes the call if communications fail.
- Named owner for every external dependency.
- Pre-start, in-progress and completion notifications defined.
- Single named point of contact during execution.
- Defined behaviour if the control room cannot be reached.
6. Rollback, recovery and hold points
A rollback plan is only credible if it is still achievable at the point of no return. The reviewer should walk the sequence and identify the step after which the original state cannot be restored within the accepted window — then confirm that a hold point sits immediately before it.
Hold points must name who acknowledges them. A hold point without an owner is a pause, not a control.
- Rollback written as a sequence, not as 'reverse the above'.
- Time required for rollback stated and compared against the accepted window.
- Point of no return identified explicitly.
- Hold point before every irreversible step, with a named acknowledger.
- Recovery actions for the credible failure modes, not only for success.
7. Testing, verification and authorisation
Testing confirms the work achieved its purpose. Verification confirms the facility is back in its intended state. They are separate checks and reviewers frequently find only the first.
Authorisation should name the approving roles and record that each of them reviewed the document at its current revision. A signature against a superseded revision invalidates the review.
| Check | Reviewer confirms |
|---|---|
| Functional test | Test proves the intended outcome, with a pass criterion |
| Return to normal | Every altered setting, mode, inhibit and isolation has a removal step |
| Redundancy proof | Restored redundancy is demonstrated, not assumed |
| Baseline comparison | Post-work readings compared against recorded pre-work readings |
| Authorisation | Correct roles, current revision, dated |
The 26-point MOP review check
Use the following as the final review register. Record Pass, Fail or Comment against every line. A failure is not cured by a general approval note; the MOP should be revised, reissued and reviewed at the new revision.
Check 1 is a readiness gate. Before working through the remaining checks, confirm that all task-specific PPE and safety equipment required by the MOP is available, fit for purpose, undamaged, within any required inspection or calibration period where applicable, and ready for use. If it is unavailable, unsuitable for the task, damaged, not fit for purpose, outside a required inspection or calibration period as required by the applicable site or task requirement, or otherwise not ready for safe use, the MOP is not ready to proceed and the review should be recorded as a fail against Check 1 pending resolution.
| # | Review check | Pass / Fail / Comment |
|---|---|---|
| 01 | Task-specific PPE and safety equipment required by the MOP is available, fit for purpose, undamaged and within required inspection or calibration periods where applicable | |
| 02 | Scope, exclusions and work boundary are unambiguous | |
| 03 | Assets use current unique references | |
| 04 | Affected upstream, downstream and shared systems are identified | |
| 05 | Required initial plant state is explicit | |
| 06 | Current drawings and references carry revisions | |
| 07 | Competent roles and supervision are named | |
| 08 | Tools, spares and calibrated test equipment are confirmed | |
| 09 | Concurrent work and degraded states are checked | |
| 10 | Operational impact and reduced resilience are stated | |
| 11 | Worst credible outcome is addressed | |
| 12 | Objective stop conditions are defined | |
| 13 | Isolation boundary and proving method are clear | |
| 14 | Each switching step names one action and one device | |
| 15 | Every critical action has an expected result | |
| 16 | Measured results have an acceptance range or criterion | |
| 17 | If-not-obtained instructions stop improvisation | |
| 18 | Dependencies have owners and confirmation points | |
| 19 | Communications cover pre-start, execution and completion | |
| 20 | Hold points precede irreversible or high-exposure steps | |
| 21 | Rollback is a timed, executable sequence | |
| 22 | Recovery covers credible failure modes | |
| 23 | Testing proves the work objective | |
| 24 | Verification proves restoration and redundancy | |
| 25 | Approvers, distribution and document revision are controlled | |
| 26 | Close-out records, defects and handover are assigned |
Approval, revision and controlled distribution
The approval block should distinguish technical review from operational authority. The technical reviewer confirms the method and engineering logic; the operations authority accepts the site state, timing and residual operational exposure. Other approvals may be required by the site's governance, but a signature should never obscure what that person is accountable for.
The issue used at the workface must match the approved revision. Withdraw superseded copies, identify the controlled execution copy, brief every executing person on changes, and ensure the control room and work party refer to the same step numbering. If the method changes after approval, stop and send the revision back through the defined review route.
8. Contingency, post-work verification and documentation
Contingency covers what happens when the work cannot be completed: overrun beyond the accepted window, a failed component that cannot be replaced tonight, or a system that will not return to its normal state. Each needs a defined safe holding condition.
Documentation closes the loop — completed procedure, readings, defects raised, asset updates and the handover entry for the next shift.
- Defined safe holding state if the work must be abandoned mid-sequence.
- Overrun escalation route with a named decision maker.
- Post-work verification signed by operations, not only by the executing team.
- Defects and observations routed to the register on the day.
- Completed procedure returned to document control at its issued revision.
The reviewer's walkthrough
Before signing, read the procedure once as the executing technician: standing in front of the plant, at night, with no access to the author. At every step ask what device is being touched, how the result is confirmed, and what happens if it does not appear.
Steps that survive that reading are executable. Steps that require the author's intent to interpret are not, however technically correct they may be.
Key takeaways
- 01Reject any scope that two competent technicians could read differently.
- 02Check affected systems past the asset itself — dependencies, controls and monitoring.
- 03Every switching step needs an action, expected result, tolerance and an if-not-obtained instruction.
- 04Identify the point of no return and put a named hold point in front of it.
- 05Testing and verification are separate checks; confirm both exist.
- 06Read the procedure as the technician before you sign it as the reviewer.
Frequently asked questions
- Who approves a data centre MOP?
- At minimum, a competent technical reviewer should confirm the method and an accountable operations authority should accept the operational state and residual exposure. Site governance may require additional approvals.
- How long before the work should a MOP be reviewed?
- Early enough to resolve comments, issue a controlled revision and brief the work party before the maintenance window. The correct lead time depends on complexity, risk and site change-control requirements.
- What are the most common reasons a MOP is rejected?
- Ambiguous scope, assumed plant state, unverifiable steps, missing hold points, weak rollback, undefined stop conditions, incomplete restoration tests and uncontrolled revision status.
- Does every task need a MOP?
- Not necessarily. The site’s work-control rules should determine when a task requires a MOP, SOP, work instruction or another control. Risk, complexity, system impact and repeatability should inform that decision.
Free resource
MOP Review Checklist
Review scope, prerequisites, step quality, hold points and recovery arrangements before approval.
Get the checklistProfessional toolkit
MOP / SOP / EOP Toolkit
Controlled templates and review tools for maintenance, standard operations and emergency response procedures.
US$59
Coming soon — not available for purchase
View product