Critical Ops HubOperations Resource
MCOPS-MNT-009Rev 1.0
Maintenance

Data Centre Maintenance Checklist

A data centre maintenance checklist is not merely a list of maintenance tasks. It is an operational control mechanism used to confirm readiness, authorisation, execution control, restoration and close-out before the activity is accepted as complete.

By Critical Ops HubPublished 10 April 2026Updated 10 June 20266 min read

Scope note

This article provides general operational guidance. Apply site-specific engineering review, risk controls, manufacturer requirements and applicable regulations before use.

Overview

The checklist does not tell a technician how to maintain the equipment. That detail belongs in the approved maintenance procedure, method of procedure (MOP), work instruction or OEM documentation. The checklist tests whether the right method and controls are in place, whether the work remains inside them and whether the facility has been returned to its intended state.

This article sets out five control stages for planned maintenance in a critical environment: planning readiness, authorisation, execution control, restoration, and close-out and evidence. If a critical check cannot be evidenced, do not treat it as complete.

Use the checklist to control the work, not describe the task

A maintenance procedure or MOP defines the technical method: the assets, sequence, isolations, tests, expected results and recovery actions. The maintenance checklist serves a different purpose. It confirms that the approved method is ready to use under the site's current conditions and that the operational controls around it remain effective from release to return to service.

A checklist built only as a task list can record that actions were ticked off while missing the conditions that make those actions acceptable. A control-structured checklist separates what must be true before work starts from what must be verified during and after it. That separation gives the operations team objective grounds to withhold authorisation, stop the activity or reject close-out.

  • Readiness checks establish that the activity can safely proceed today.
  • Authorisation checks establish who accepted the residual risk.
  • Execution checks establish that the work stayed inside the approved method.
  • Restoration checks establish that the plant is back in its intended state.
  • Close-out checks establish the record.

Define the maintenance activity before applying the checklist

A single checklist can control the work, but it cannot make every maintenance activity equivalent. Inspection observes condition without necessarily intervening. Preventive maintenance performs planned tasks intended to manage known degradation. Corrective maintenance restores a defect or failed function. Statutory or compliance activity satisfies a specific legal, regulatory or authority requirement. OEM activity follows the equipment manufacturer's instructions, limits or warranty conditions. One job can carry more than one classification, but the basis for each requirement should remain visible.

The classification matters because it changes the trigger, competence, evidence and acceptance criteria. A visual inspection may close with an observation record; intrusive preventive work may require isolation, post-work testing and a controlled maintenance window; corrective work may start from a defect and finish only when both function and redundancy are restored. Compliance and OEM requirements should be traced to the applicable source rather than rewritten as an unsupported universal interval.

Do not apply one generic monthly, quarterly or annual frequency across the asset base. Set the interval from the combination of applicable requirements, manufacturer information, asset function, failure mode, duty, condition, operating environment and site risk assessment. Where those sources conflict, resolve the conflict through the site's engineering and governance process before scheduling the work.

Maintenance activity distinctions
ActivityPrimary purposeTypical triggerCompletion evidence
InspectionObserve condition or statusRoute, event, condition or review planRecorded condition, readings, defects and escalation
Preventive maintenanceManage an identified degradation mechanismTime, usage, condition or risk-based planCompleted task, measurements, acceptance criteria and findings
Corrective maintenanceRestore a defect or lost functionDefect, alarm, failed test or breakdownFunction test, restored configuration and defect close-out
Statutory / complianceMeet an applicable external obligationRequirement set by the relevant jurisdiction or authoritySpecified certificate, test record or controlled report
OEM requirementMaintain equipment within manufacturer instructions or support termsCurrent manufacturer documentationOEM-prescribed record, reading, part or test result

Stage 1 — Planning readiness

Readiness is assessed against the state of the facility on the day, not against the state assumed when the activity was scheduled. A planned maintenance activity that was low risk at the point of scheduling can become unacceptable because a parallel system went into alarm overnight.

Planning readiness checks
CheckFrequencyResponsible partyEvidence
Approved procedure issued at the current revisionEvery activityMaintenance plannerProcedure reference and revision number
Redundancy state of the affected system confirmedEvery activityShift leadCurrent plant state record
Concurrent works checked for conflictEvery activityMaintenance plannerWork schedule review
Spares, consumables and test equipment confirmed availableEvery activityMaintenance supervisorParts and calibration records
Personnel competency and site induction confirmedEvery activityMaintenance supervisorCompetency record
Client or stakeholder notification issued where requiredWhere impact is possibleOperations managerNotification record

Apply discipline-specific checks

The control stages are common, but the technical questions must follow the system being maintained. The table below is a review aid, not a substitute for the approved procedure, current drawings, OEM information or the competence required for the work. Mark each line Pass, Fail or Not Applicable and attach the evidence reference; a blank result should not be treated as acceptance.

Discipline checklist — add Pass / Fail / N-A and evidence in the working record
SystemPractical checksResult / evidence
Utility, HV/LV and switchboardsConfirm source configuration, protection status, isolation boundary, switching authority, arc-flash and access controls where applicable, and the available alternate path.Pass / Fail / N-A; switching sheet and readings
UPS and battery systemsConfirm load and redundancy, bypass availability, battery hazards, monitoring state, expected alarms, firmware or settings controls, and the post-work functional test.Pass / Fail / N-A; test report and alarm log
Generators and fuel systemsConfirm start availability of remaining sets, fuel and starting systems, synchronising or transfer constraints, emissions or test restrictions where applicable, load-test method and cooldown/restoration.Pass / Fail / N-A; run record and readings
Power distributionTrace affected boards, PDUs and downstream loads; confirm labelling, phase/load implications, protective-device state and restoration of all temporary supplies.Pass / Fail / N-A; marked drawing and load record
Cooling and mechanical plantConfirm duty/standby state, thermal ride-through, valves and controls affected, drainage or leak controls, environmental trend points and recovery criteria.Pass / Fail / N-A; trend and leak check
BMS, EPMS and controlsIdentify points, alarms, inhibits, overrides, setpoints and integrations affected; back up settings where required and verify automatic control after work.Pass / Fail / N-A; point list and post-work trend
Fire and life safetyConfirm the approved impairment process, affected zones and interfaces, compensating controls, notifications and independent removal of every isolate or inhibit.Pass / Fail / N-A; impairment and reinstatement record
Security and accessConfirm controlled access, escort needs, door or alarm impairments, keys or credentials issued, work-area security and restoration of monitoring.Pass / Fail / N-A; access and alarm record

Control contractors, records and escalation

Specialist knowledge does not transfer operational authority to a contractor. The site operations team retains control of access, plant state, approved procedures, permits and isolations, hold points, communications, escalation and return-to-service acceptance. Before mobilisation, confirm the contractor's scope, competence evidence, supervision, tools, calibrated test equipment, spares, method, expected alarms and waste or housekeeping controls. At the pre-start briefing, make the work boundary, communications route and stop-work authority explicit.

Maintenance records should identify the asset, task basis, procedure revision, personnel, start and finish times, readings, parts used, defects, temporary conditions and acceptance evidence. When a stop condition is reached, hold the activity in the defined safe state and use the approved escalation route. Do not change the method informally at the plant.

  • Stop if a prerequisite fails or cannot be evidenced.
  • Stop if the actual plant state does not match the approved starting state.
  • Stop if the required redundancy is unavailable.
  • Stop if the asset, device reference or work boundary is wrong or unclear.
  • Stop if a required test fails or its acceptance criterion is not met.
  • Stop and assess any unexpected alarm or indication.
  • Stop if an isolation cannot be applied and proven as approved.
  • Escalate before the accepted maintenance window or reduced-state duration is exceeded.
  • Do not close the work with a temporary condition that cannot be removed or formally controlled.
  • Do not return the system to service until restoration is demonstrated.
  • Stop and return any method change through the approved review and authorisation route.

Stage 2 — Authorisation

Authorisation is the point at which a named person accepts that the facility will operate in a reduced or altered state for a defined period. It should be a discrete step with a time, a name and a stated duration — not an implied consequence of the work appearing on a schedule.

A worked authorisation sequence for a single UPS module inspection illustrates the intent: confirm the redundant module is healthy and not itself under a deferred defect; confirm no other power-path work is in progress; confirm the maintenance bypass has been proven within its required interval; record the accepted duration; name the person who can stop the work; and log the time the reduced state began.

  • Named authoriser and named person with stop authority.
  • Stated start time and maximum accepted duration for the reduced state.
  • Conditions that automatically end the authorisation.
  • Escalation contact if any condition is breached.

Stage 3 — Execution control

During execution, the checklist's job is to keep the activity inside the approved method. The most useful execution checks are the ones that confirm state changes were made as written, and that any deviation was stopped rather than improvised.

Execution controls
CheckWhenResponsible partyEvidence
Isolation applied and proven against the approved sequenceBefore first interventionExecuting technicianSigned isolation record
Hold point acknowledged by operations before the critical stepAt each hold pointShift leadHold point sign-off
Deviation stopped and referred rather than resolved in the fieldOn any deviationExecuting technicianDeviation record
Alarms generated by the work identified and confirmed expectedContinuousControl roomAlarm log annotation
Duration tracked against the accepted windowContinuousShift leadTime record

Stage 4 — Restoration

Restoration is where reduced states become permanent if the checklist is weak. Every temporary condition introduced during the work needs an explicit removal check, and the system's normal operating state needs to be confirmed positively rather than assumed from the absence of alarms.

  • All temporary isolations, bypasses, inhibits and overrides removed and individually confirmed.
  • Control settings, setpoints and operating modes returned to the recorded normal state.
  • Redundancy proven — not inferred — by observing the restored system carrying or capable of carrying its duty.
  • Alarm state returned to the pre-work baseline, with any remaining alarms explained.
  • Reduced-state period formally closed with a time and a name.

Stage 5 — Close-out and evidence

Close-out converts the activity into information other people can act on: the asset record, the defect list, the next planned intervention and the operational handover. A checklist that ends at the last technical step loses all of it.

Close-out record
RecordDestinationResponsible party
Completed procedure with signatures and timesDocument controlMaintenance supervisor
Defects and observations raisedDefect registerExecuting technician
Asset condition and readingsAsset register / CMMSMaintenance planner
Parts consumed and spares to reorderSpares registerMaintenance supervisor
Operational summary for the next shiftShift handover recordShift lead

Keep the checklist inside its boundary

The five stages stay the same across disciplines, but the evidence changes. Electrical work may depend on isolation proving and redundancy confirmation. Mechanical work may depend on sequencing and thermal recovery—a cooling system restored electrically is not restored operationally until the required environmental condition is demonstrated. Fire and life-safety work may depend on controlled impairments and independent confirmation that every isolate or inhibit has been removed.

The checklist does not set maintenance strategy, select task intervals, create the maintenance plan or define the maintenance window. It also does not replace the procedure or MOP that controls the technical sequence. Those subjects remain with their dedicated guidance. The free Data Centre Maintenance Checklist provides the working record for these five stages; the Data Centre Maintenance Playbook covers the wider management system around maintenance planning, governance, procedures, asset information, risk controls, contractor management, performance and continuous improvement.

Key takeaways

  • 01Assess readiness against the facility's state today, not the state assumed at scheduling.
  • 02Make authorisation a discrete, named, time-bound acceptance of a reduced state.
  • 03Give every critical step a hold point that operations must acknowledge.
  • 04Stop deviations and refer them; never resolve a method change in the field.
  • 05Prove restoration positively — removed bypasses, confirmed setpoints, demonstrated redundancy.
  • 06Close out into the records other people depend on: defects, assets, spares and handover.

Frequently asked questions

How often should data centre maintenance be carried out?
There is no single interval for every asset. Use applicable statutory requirements, current OEM instructions, site requirements, duty, condition, failure modes and risk assessment to set and review each task frequency.
What is included in a data centre maintenance checklist?
Planning readiness, work classification, authorisation, discipline-specific technical checks, isolation and execution controls, restoration, records, defects and escalation. Each check should identify its owner and evidence.
Who is responsible for data centre maintenance?
Accountability is shared but must be explicit: the asset or maintenance owner plans the task, competent personnel execute it, and the operations authority controls plant state, authorisation, hold points and return to service.
What is the difference between preventive and corrective maintenance?
Preventive maintenance is planned to manage an identified degradation mechanism before functional failure. Corrective maintenance responds to a defect or loss of function and must restore both the asset and the required operating configuration.

Free resource

Data Centre Maintenance Checklist

Plan, control and close maintenance activities with clear readiness and restoration checks.

Get the checklist

Professional toolkit

Data Centre Maintenance Playbook

A practical operating framework for maintenance strategy, planning, controlled execution and performance review.

US$49

Coming soon — not available for purchase

View product

Continue reading

Related articles

Back to the Maintenance pillar →