1. Normal-source outage
Time during which the normal utility or facility source is unavailable to serve its intended role—regardless of whether alternate paths, storage or generation maintain the load.
Data-center reliability knowledge
What does an electrical outage actually cost an AI data center? The answer depends on which clock you are timing—and on what the workload was doing when the event occurred.
A prolonged utility outage may produce no compute interruption if alternate sources and stored energy perform correctly. Conversely, a brief rack-power interruption can stop a training workload and cause a much longer recovery process than the electrical gap itself.
Illustrative article. Timing values, sequences and cost examples are modeling assumptions for explanation. They are not NVIDIA, OEM or site-specific specifications unless explicitly cited as published references below.
Why one duration is not enough
Operators, insurers, utilities and workload owners often speak about “the outage” as if it were a single interval. In practice, several distinct durations may coexist in one event—and each may drive a different part of the cost.
Service restored ≠ full resilience restored. A facility may resume revenue-bearing work while still operating with reduced margin, partially recharged batteries, temporary generation or unrepaired equipment.
Part 1
The following eleven clocks are frequently conflated in conversation and reporting. Each should be defined explicitly before comparing facilities, contracts or mitigation options. Event durations span many orders of magnitude—from milliseconds at the rack to hours or days for full restoration—consistent with the POI time-scale discussion on Load Behavior & POI.
Time during which the normal utility or facility source is unavailable to serve its intended role—regardless of whether alternate paths, storage or generation maintain the load.
Time from fault inception through detection, primary protection and, when required, breaker-failure or backup protection. This interval shapes whether downstream ride-through is challenged immediately or after a brief disturbance.
The brief period during which server PSU capacitors maintain internal DC rails after rack input power is lost. Published ATX-family design guidance specifies hold-up on the order of milliseconds to tens of milliseconds depending on loading and product class—see references. An illustrative 10 ms modeling value is used in examples below; it is not a universal server or rack specification and may differ from cited ATX minima and from OEM designs. Actual values must come from the server/OEM specification for the equipment in service.
Time from loss of rack input until rack-level storage accepts the load. Distinguish:
The critical design relationship at the rack is:
storage takeover time < PSU hold-up time
If takeover exceeds hold-up, the server may reset even though facility BESS or another upstream source exists but has not yet reached the rack.
Time for which stored energy can sustain the rack at the required load. Rack BBU autonomy and facility BESS endurance differ in energy, power, state of charge and recharge requirements. 10 minutes appears below only as an illustrative rack-bridging scenario—not as an NVIDIA or universal data-center specification. Autonomy depends on state of charge, load, temperature, battery health and design margin.
Separate sub-intervals often matter:
“Generator started” is not the same as “the required rack load has been successfully transferred.”
Time during which acceptable electrical power is unavailable at the rack after considering all normal feeds, alternate paths, storage and PSU hold-up.
Time during which servers or accelerators cannot remain operational because their internal power requirements are not maintained.
Time until the application or workload is again performing its intended function. This can include:
For illustration only, consider a training job with a 20-minute checkpoint interval (configuration- and workload-dependent). A failure may lose between zero and nearly one full interval of work. If failures occur uniformly between checkpoints, the expected lost progress is approximately 10 minutes under that illustrative assumption—not a universal training loss rule.
progress recovery time = useful-compute restoration time + time required to repeat lost work
This measure is important for training and long-running batch jobs but is generally not applicable to interactive inference in the same way—there, lost requests rather than repeated gradient steps dominate.
Time until:
Terminology
Colloquial outage language often mixes layers. The table maps common terms to the clocks defined above.
| Colloquial term | Primary clock(s) | Notes |
|---|---|---|
| Utility outage / utility-supply interruption | 1. Normal-source outage | Utility unavailable in its normal role; rack power may continue on BESS, generator or alternate feed. |
| Normal-source outage | 1. Normal-source outage | Same as intended utility or designated normal facility source—not necessarily loss of all rack input. |
| Rack-power interruption | 7. Rack-power interruption | Loss of acceptable electrical input at the rack after hold-up and bridging. |
| IT interruption | 7–9 | Often used loosely; may mean rack power, hardware down or service unavailable. Specify which layer is meant. |
| Compute interruption / compute-hardware interruption | 8. Compute-hardware interruption | Servers or accelerators cannot remain powered and operational. |
| Service interruption / useful-compute interruption | 9. Useful-compute or service interruption | Application or workload not performing its intended function; may persist after rack power returns. |
| Progress recovery / computational recovery | 10. Pre-fault computational progress recovery | Includes restart plus repetition of work lost since the last valid checkpoint (training/batch context). |
| Full restoration / resilience restoration | 11. Full-resilience restoration | Normal configuration, redundancy, recharge and reserve margin—not the same as service restored. |
InfraRel-compatible consequence measures related to these clocks include expected unserved energy (electrical energy not delivered), unserved computation (useful work not completed) and lost accelerator-hours (compute capacity-time lost to interruption and recovery).
Worked example
Illustrative sequence only. Actual timing depends on protection, switching, storage state and load.
Duration summary. Normal-source outage has a positive duration. Rack-power interruption, compute-hardware interruption and useful-compute interruption durations are zero in this illustrative case. Operational cost may still include fuel, battery cycling, operator response and reduced reserve margin until full resilience returns.
Worked example
The same initiating electrical event can produce multiple non-zero clocks and a long recovery tail.
Electrical service may return to the rack at 20 minutes from backup while the utility normal source remains out until later. Useful compute may remain unavailable until ~35 minutes and pre-fault progress may not be restored until ~42 minutes in this illustration. Directly attributable SLA credits, repeated-work cost and lost accelerator-hours attach to those longer intervals—not to the utility-outage clock alone.
These values are scenario assumptions for teaching. They do not represent every data center, every checkpoint policy or every backup design.
Worked example
Illustrative sequence showing that the normal source may remain available while a brief electrical disturbance still stops useful computation for much longer.
Duration summary. Normal-source outage ≈ 0. Rack-power interruption is very short. Useful-compute interruption and progress-recovery duration can dominate cost through idle accelerator-hours, unserved computation and lost training progress—despite minimal utility-visible outage time.
Visual summary
The figure shows a subset of clocks from Example B—not all eleven—prioritizing readability. Rack power may return from backup while the utility normal source remains unavailable.
Boundary
Read this section before interpreting the cost categories below. This article estimates operational-continuity costs attributable to an electrical initiating event and the resulting sequence of protection, switching, backup-power, rack-power, compute-interruption and workload-recovery states.
A cyber event may initiate an electrical outage, and an electrical outage may expose cyber or data-recovery weaknesses. This article calculates the operational consequences of the electrical and computing interruption. Cybersecurity, AI-behavioral, misuse and societal consequences require additional risk models and are not included in the cost estimate.
Do not describe every excluded cost as “indirect.” Some cyber and AI incidents create direct financial costs. They are separate risk domains outside the defined electrical-outage cost boundary.
Authoritative frameworks for adjacent domains include the NIST Cybersecurity Framework 2.0, the NIST AI Risk Management Framework, and relevant critical-infrastructure cybersecurity guidance. This article does not provide a complete enterprise-risk, cyber-risk or societal-cost assessment.
Part 2
Outage cost cannot be calculated from electrical outage duration alone. The categories below are bounded by Scope of the cost assessment. The same utility interruption may be inexpensive if bridged successfully, or very expensive if it causes checkpoint rollback, unserved computation and lost accelerator-hours.
Total outage cost = lost service + idle compute + repeated work + recovery + SLA consequences + operational intervention + equipment/data consequences + reserve restoration
Generalized reputational impact, litigation, regulatory penalties and enterprise-valuation effects are out of scope unless separately modeled with explicit assumptions.
Workload comparison
An interruption can create checkpoint rollback, repeated work, lengthy restart, cluster reassembly, idle accelerator-hours and delayed model-development schedule. Training interruptions can have a long recovery tail and may be especially costly per affected job.
An interruption can cause immediate user impact, failed or delayed requests, traffic rerouting and defined SLA credits. Inference outages may be shorter in recovery but can immediately affect very large numbers of users.
Cost depends on workload, scale, traffic, checkpoint policy, recovery automation, redundancy and contractual commitments. Training is not always costlier than inference, and inference is not always cheaper—each event must be evaluated in context.
Qualitative comparison
The table below is qualitative and configuration-dependent. It illustrates why electrical duration alone is an unreliable proxy for total consequence.
| Electrical event | Electrical duration | Compute consequence | Potential cost |
|---|---|---|---|
| Utility loss successfully bridged | Long | None | Low operational/energy cost |
| Brief rack interruption causing server reset | Very short | Long workload recovery | Potentially high |
| Cooling loss with rack power intact | No rack-power outage | Throttling or shutdown | Potentially high |
| Network isolation with healthy compute | No electrical outage | Service unavailable | Potentially high |
| Transfer failure / failed load transfer | Short or none at utility | Rack or facility load not accepted | Potentially high |
| Control-power loss | May be brief or none | Protection, transfer or cooling logic unavailable | Potentially high |
| Generator started but transfer failed | Backup running | Rack still unserved; long recovery possible | Potentially high |
| Training-job interruption | Short or moderate | Checkpoint rollback and repeated work | Workload-dependent |
| Inference-service interruption | Short or moderate | Immediate user/SLA impact | Traffic-dependent |
Qualitative illustration only. Actual cost requires site-specific inputs and explicit assumptions.
Electrical: Long · Compute: None · Cost: Low operational/energy
Electrical: Very short · Compute: Long recovery · Cost: Potentially high
Electrical: No rack-power outage · Compute: Throttling/shutdown · Cost: Potentially high
Electrical: None · Compute: Service unavailable · Cost: Potentially high
Electrical: Short/none at utility · Compute: Load not accepted · Cost: Potentially high
Electrical: Brief/none · Compute: Logic unavailable · Cost: Potentially high
Electrical: Backup running · Compute: Rack unserved · Cost: Potentially high
Electrical: Short/moderate · Compute: Rollback/repeat · Cost: Workload-dependent
Electrical: Short/moderate · Compute: User/SLA impact · Cost: Traffic-dependent
Inputs
Missing inputs should be represented as ranges, scenarios or explicitly marked unknowns—not silently invented. InfraRel-style studies may report expected unserved energy, unserved computation and lost accelerator-hours alongside dollar-valued outage cost.
Emerging methodology
General Reliability is extending InfraRel to connect the complete chain from network configuration through outage cost—under development and not yet a site-specific validated data-center assessment.
Network configuration → Fault and protection response → Time-sequenced switching → Ride-through → Rack-power consequence → Workload recovery → Outage cost
We welcome collaboration with data-center operators, utilities, electrical and cooling OEMs, server and storage specialists, network and workload engineers, insurers and risk specialists, and researchers who can contribute operational data, validation experience and domain requirements.
Sources
Sources are grouped by type. Marketing language is not treated as established engineering fact.
Design-guide hold-up requirements for desktop/platform PSUs. Server and rack PSUs may differ; use OEM data for in-service equipment.
Illustrates that transfer intervals are product- and mode-dependent and may exceed server hold-up unless double-conversion or matched designs are used.
Public documentation on checkpoint save/load behavior relevant to training recovery after interruption.
Survey-based perspective on outage frequency, severity and reported cost magnitudes. Supports general discussion of outage-cost scale—not this article’s electrical-to-computational clock taxonomy. Full report requires membership.
Facility-level reference material on power-block architecture, stored energy and ride-through concepts. Reference designs are illustrative, not site-specific specifications.
Evidence classes used in this article: published specification/documentation; general industry reference; illustrative assumption (timing scenarios); owner/OEM-required input (site-specific switching, storage, workload and cost data).