, ,
1–2 minutes

Single-feed failure in a data center: the first 60 minutes that matter

In a 2N A/B electrical design, the first hour is about confirmation, isolation, and load stability. The goal is to prove that the critical IT load stayed protected, that the surviving path is not near a limit, and that the failed side can be investigated without creating a second event.

Minutes 0-10: verify and classify

  • Check BMS/EPMS alarms and identify whether the event is upstream utility, switchgear, UPS, PDU, or downstream distribution.
  • Confirm dual-corded systems stayed online and ATS/STS devices transferred automatically.
  • Validate the generator and UPS state on the failed path before any switching action.

Minutes 10-30: protect the live path

  • Inspect rack PDU loading on the surviving feed and compare it with remaining headroom.
  • Identify any single-corded devices that dropped offline and log them separately.
  • Check inlet temperatures, CRAC/CRAH response, and any cooling alarms that appear after the transfer.

Minutes 30-60: isolate and prepare recovery

  • Keep the failed path isolated while the electrical team traces the fault.
  • Stage spare parts and review any BMC or controller alarms triggered during the event.
  • Capture timestamps, alarm names, and operator actions for the incident log.
  • Prepare a short leadership update that states what failed, what remained stable, and what the next decision point is.

This operating model matches the direction in the Uptime Institute Annual Outage Analysis 2025, which continues to show power as a leading cause of impactful outages. It also aligns with Schneider Electric’s UPS system design guidance, DOE guidance on efficient UPS systems, and the DOE best practices guide for data center design.