How to Write a Data Centre MOP
A method of procedure is the written sequence that lets planned work happen on live infrastructure without relying on the author being present. It is written for the person executing it at their least alert hour, not for the person approving it.
Scope note
This article provides general operational guidance. Apply site-specific engineering review, risk controls, manufacturer requirements and applicable regulations before use.
Overview
This article walks the practical process end to end: defining scope, identifying affected systems, setting prerequisites, assessing operational risk, writing the sequence, placing hold points, defining testing and verification, writing rollback and recovery, planning communications, review, approval, execution and close-out.
It ends with a complete short MOP for a routine activity, written out in full, so the structure is visible rather than described.
Separate the five control layers
A usable MOP is more than a technical sequence. Write and review five connected layers. The technical sequence states what is done to the equipment. Operational controls state the plant conditions that must remain true. Approval controls state who accepts the work and at which revision. Execution controls govern holds, communications, records and deviations. Recovery controls define how the team returns to a known state when the expected result is not obtained.
Keeping these layers visible prevents a common drafting error: placing every safeguard in the risk assessment while leaving the execution steps silent. A technician should not need to cross-reference an abstract control to discover that operations must acknowledge a hold point before the next switch is operated.
| Layer | Question it must answer | Typical evidence |
|---|---|---|
| Technical sequence | What action is performed, on which device, and what result is expected? | Numbered steps and readings |
| Operational controls | What plant state, redundancy and environmental conditions must be maintained? | Prerequisites, limits and stop conditions |
| Approval controls | Who reviewed and accepted this exact revision? | Approval block and revision record |
| Execution controls | Who communicates, holds, records and stops the work? | Briefing, hold-point and step sign-offs |
| Recovery controls | How is a safe known state reached if the plan cannot continue? | Rollback and failure-mode recovery sequences |
Step 1 — Define the scope
Write the scope as one sentence that a competent technician can only read one way, then add the exclusions. If the scope needs a paragraph, the activity is probably two MOPs.
Name assets by their register reference. 'The standby generator' is not a scope; 'Generator GEN-02, register reference AST-0412' is.
Gather the controlled inputs before writing
Do not draft from memory or from an old MOP with the asset names changed. Gather the current single-line diagrams, schematics, asset register references, cause-and-effect information, control narratives, OEM instructions, existing operating procedures, defect status, maintenance history relevant to the task and the site's work-control requirements. Confirm which source is authoritative and record its revision.
Walk the intended work boundary where practical with operations and the executing specialist. Compare labels and device references with the drawings. Identify temporary arrangements, open defects and abnormal configurations early; they may invalidate the assumed sequence before writing begins.
Step 2 — Identify affected systems
Work outwards from the asset: what it feeds, what feeds it, what shares its control system, what monitors it, and what interlocks with it. Fire detection, BMS points and cooling that depends on the same power path are the ones most often missed.
Record for each affected system whether it is unaffected, degraded or unavailable during the work, and for how long.
Step 3 — Set the prerequisites
Prerequisites are verifiable on the day, each with a pass criterion. Write them so a shift lead can refuse the start without needing to argue technical judgement: the redundant unit is healthy, no concurrent works on the same path, drawings at the stated revision are on site, test equipment in calibration, competent personnel present, notifications issued.
Step 4 — Assess operational risk
State the worst credible operational outcome and the duration of any reduced-resilience state. Then define the stop conditions objectively — a second failure on the same path, an environmental reading outside its band, a duration overrun, a missing prerequisite discovered mid-sequence.
Step 5 — Write the sequence
One action per step, one device per step, in execution order. Each step carries the same five elements, and a step missing any of them will be rejected at review.
| Element | Example |
|---|---|
| Action | Open input isolator IS-2A at UPS-2 |
| Method | Local operation by authorised electrical person |
| Expected result | Isolator indicates open; UPS-2 shows input loss |
| Tolerance | Downstream board voltage remains within its recorded normal band |
| If not obtained | Stop. Do not proceed to step 6. Contact the shift lead. |
Step 6 — Place the hold points
Put a hold point before every irreversible step and before every step that removes the last layer of redundancy. Name the role that acknowledges each one. The hold point should state what the acknowledger is confirming, not simply that they agree to continue.
Step 7 — Define testing and verification
Testing proves the work achieved its purpose. Verification proves the facility is back in its intended state. Write both, with pass criteria, and record pre-work readings so the post-work comparison has a baseline.
Step 8 — Write rollback and recovery
Write rollback as its own numbered sequence with its own expected results. State how long it takes and identify the point after which it is no longer available within the accepted window. Add recovery actions for the credible failure modes — a breaker that will not reclose, a unit that will not restart, a control system that does not return to automatic.
Step 9 — Communications, review and approval
Define who is notified before, during and after, and the single point of contact during execution. Then submit the document for review against the MOP review checklist, and record approval by role, name, date and document revision.
Approval against a superseded revision is the most common document-control failure. Issue the revision that will actually be executed.
Review the MOP as an executable control
Use separate technical and operational reviews. The technical reviewer challenges device selection, sequence, interlocks, testing and recovery. The operations reviewer challenges timing, redundancy, concurrent risk, alarm response, communications, authority and restoration. Resolve comments in the document; do not leave essential controls in email threads or meeting notes.
After revision, perform a step-by-step tabletop walkthrough with the people who will execute and control the work. Read device references aloud, identify who acts and who observes, and confirm where signatures or readings are recorded. Any step that depends on unwritten local knowledge is not ready for issue.
Step 10 — Execution and close-out
During execution the document is the record: sign each step, record readings, annotate any deviation and stop rather than adapt. At close-out, return the completed procedure to document control, raise defects and observations, update the asset record, and write the handover entry for the incoming shift.
Afterwards, feed what was learned back into the document. A MOP that was hard to follow in the field should be revised before it is used again.
A complete short MOP
The following is a full, deliberately small MOP for a routine filter changeout on a computer room air handling unit in an N+1 cooling arrangement. It is illustrative and must be adapted to site conditions before use.
Scope: Replace the return-air filter set on CRAH-03 (AST-0214), Data Hall 1. Excludes any work on the chilled water circuit, controls or electrical supply.
Affected systems: CRAH-03 unavailable for the duration. Data Hall 1 cooling operates at N during the activity. BMS alarms expected for CRAH-03 airflow and unit status.
Prerequisites: CRAH-01, 02 and 04 healthy and carrying load; hall temperature within its normal band; no concurrent mechanical works in Data Hall 1; replacement filters on site and correct part number verified; two competent technicians; operations notified.
Risk and stop conditions: Worst credible outcome is a rise in hall temperature if a second unit fails while CRAH-03 is out. Accepted duration 60 minutes. Stop if hall temperature rises beyond its normal band, or if any other CRAH alarms.
Sequence: (1) Record hall temperature and the status of all four CRAH units. (2) HOLD POINT — shift lead confirms N+1 is intact and accepts the reduced state, recording the start time. (3) Place CRAH-03 into maintenance mode at the BMS; confirm the unit reports maintenance mode. (4) Stop CRAH-03 locally; confirm fan stopped. (5) Remove and replace the filter set; record the part number fitted. (6) Close panels and confirm secure. (7) Restart CRAH-03 locally; confirm fan running and airflow established. (8) Return CRAH-03 to automatic at the BMS; confirm the unit reports automatic. (9) Confirm all CRAH-03 alarms cleared. (10) Record hall temperature and confirm it is within its normal band.
Rollback: If the filter set cannot be fitted, refit the original filters, restart CRAH-03 and return it to automatic. Rollback time approximately 15 minutes; available up to step 6.
Verification and close-out: Post-work readings compared against step 1. Shift lead signs the reduced state closed. Completed procedure to document control, filter change recorded against AST-0214, handover entry written.
Key takeaways
- 01Write for the technician executing at their least alert hour, not for the approver.
- 02Name every asset by its register reference; one action and one device per step.
- 03Every step needs an expected result, a tolerance where measurable, and an if-not-obtained instruction.
- 04Hold points go before every irreversible step and every loss of the last redundancy layer.
- 05Rollback is its own sequence with its own duration and its own expiry point.
- 06Approve the revision that will actually be executed, and close out into the asset and handover records.
Frequently asked questions
- What is the standard structure of a MOP?
- A practical structure covers document control, scope and exclusions, references, affected systems, prerequisites, operational risk, roles and communications, numbered steps, hold points, testing, verification, rollback, recovery, approvals and close-out.
- How long should a MOP be?
- Long enough to control the activity without hiding critical actions in unnecessary text. Complexity, system impact and recovery needs determine length; clarity and executability matter more than page count.
- What is a hold point?
- A mandatory pause before a defined step where a named role confirms stated conditions and authorises continuation. It should record who acknowledged it, when, and what was verified.
- What is the difference between a MOP and a work instruction?
- A MOP controls a specific planned activity and its operational exposure from start to close-out. A work instruction usually explains how to perform a particular technical task; it may be referenced by a MOP but does not automatically provide the wider operational controls.
Free resource
MOP Review Checklist
Review scope, prerequisites, step quality, hold points and recovery arrangements before approval.
Get the checklistProfessional toolkit
MOP / SOP / EOP Toolkit
Controlled templates and review tools for maintenance, standard operations and emergency response procedures.
US$59
Coming soon — not available for purchase
View product