Once the critical load is stable, the floor team should spend the next hour on verification rather than recovery shortcuts. This is the stage where disciplined observation matters more than movement.
Floor checks
- Review rack PDU loading and breaker status on the active side.
- Identify single-corded devices that lost power and record them as a separate action list.
- Check CRAC/CRAH response, inlet temperatures, and any hot-spot alarms.
- Verify generator fuel level, UPS alarms, and transfer behavior.
- Pull BMC or controller alerts from affected racks to see whether the event triggered secondary faults.
How the team should work
- Assign one person to the electrical path, one to cooling, one to IT alarms, and one to the incident log.
- Compare actual load to remaining headroom before any restoration step is approved.
- Keep the failed path isolated until the fault is known and the restoration method is agreed.
For the monitoring side, the DOE’s Energy-Efficient Cooling Control Systems for Data Centers page is useful because it covers real-time monitoring, visualization, analysis, and feedback control. The DOE best practices guide and the Uptime Institute outage analysis are also useful references when you are explaining why you want data before action.