← Tools & Methods · Data Centers hub →

Reliability Analysis Platform

InfraRel: From Component Failures to Infrastructure Consequences

InfraRel is an emerging reliability-analysis platform for evaluating interconnected electrical and supporting infrastructure systems. It extends the foundation of SUBREL, which has been used in power-industry applications to evaluate and compare alternative substation configurations.

Under active development

Why InfraRel

Reliability Does Not Stop at the Electrical Connection

Critical infrastructure depends on more than the availability of utility power. A data center, industrial facility or other essential service may also depend on local generation, batteries, switching systems, cooling, communications, controls, buildings and operational processes. A failure in any required subsystem can interrupt the service even when other parts of the system remain available.

InfraRel expands the modeled boundary so that failures, dependencies, restoration actions and their consequences can be evaluated across the complete service-delivery chain.

  1. Grid and local sources
  2. Facility power
  3. Cooling, communications and controls
  4. Critical load
  5. Useful service

Power is also required by cooling, communications, controls, offices and supporting facilities — so electrical availability and supporting-system availability are coupled.

Foundation

A New Extension Built on an Established Foundation

InfraRel is relatively new, but its underlying reliability methodology is not. It extends SUBREL, a reliability-analysis program developed and used in power-industry applications to compare alternative substation configurations.

The established foundation includes the representation of component outages, system configurations, switching and restoration actions, and their effects on continuity of supply. InfraRel extends this approach to interconnected infrastructure systems and broader operational consequences.

SUBREL foundation

  • Substation configuration comparison
  • Component-outage representation
  • Switching and restoration actions
  • Continuity-of-supply indices
  • Power-industry applications

SUBREL overview →

InfraRel extension

  • Electrical and supporting infrastructure
  • Cross-system dependencies
  • Critical-load and service consequences
  • Workload restoration and lost computation
  • Outage-cost evaluation
  • Scenario and parameter-matrix studies

InfraRel is emerging; it does not yet share SUBREL’s application history.

System boundary

An Integrated System Boundary

Utility supply and grid connections

External supply paths and interconnection points that feed the facility.

Local generation

On-site generation that can support ride-through, islanded or backup operation.

Batteries, UPS and stored energy

Short-term continuity and bridging between sources.

Automatic and manual transfer arrangements

Transfer paths that change the active supply configuration.

Switchgear and protection

Isolation, protection and switching that shape outage and restoration sequences.

Cooling and thermal support

Thermal systems required to keep critical loads operable.

Communications and network systems

Networks that support monitoring, control and operational continuity.

Monitoring and control systems

Supervisory and control functions that govern response and recovery.

Buildings, offices and supporting services

Supporting facilities that may be required for continued operation.

Critical computational or operational loads

The loads whose interruption defines service failure for the study.

The model can be simplified to match the study objective. Not every system or component must be represented at the same level of detail, but material dependencies and exclusions should be identified explicitly.

Methodology

From Failures to Service Consequences

  1. Define configurations, operating states and system boundaries.
  2. Represent components, connections and required dependencies.
  3. Generate applicable outage events and failure combinations.
  4. Evaluate switching, isolation and restoration actions.
  5. Determine interrupted loads, services and restoration durations.
  6. Calculate reliability, consequence and economic indices.

The objective is not merely to determine whether power is restored. For a data-center application, the relevant endpoint is whether useful computation continues or is restored within an acceptable time and cost.

Indices & consequences

Beyond a String of Nines

A single availability percentage is insufficient for many infrastructure decisions. InfraRel organizes complementary measures so frequency, duration, operational impact and cost can be examined separately.

Expected interruption frequency

How often disruptive events are expected to occur.

Physical interruption duration

How long physical supply or support is interrupted.

Workload interruption and recovery duration

How long useful work is lost, including restart and recovery.

Expected unserved energy

Energy not delivered because of interruptions.

Expected unserved computation

Useful computational work not completed because of interruptions.

Lost accelerator-hours

Accelerator or compute capacity-time lost to disruption and recovery.

Restoration and restart consequences

Operational effects of switching, restart and return to service.

Expected outage cost

Economic consequences associated with interruptions and recovery.

Two systems with the same availability can experience very different interruption frequencies, restoration times, lost computation and financial consequences. InfraRel therefore separates the frequency of disruptive events from their physical and operational duration.

Related reading: outage duration clocks and cost boundaries →

Related methodology: Data-Center Power Architecture & Reliability → Configuration, switching and common-mode structure on that page can inform InfraRel case definition where scoped.

Related methodology: Grid Capacity & Flexibility Assurance Framework → InfraRel may support configuration and consequence analysis within that broader assurance workflow where scoped.

Related methodology: Load Behavior & POI → Operating-state, ramp, rebound and model-requirement framing can inform InfraRel scenarios where scoped.

Uncertainty & ranges

From Individual Cases to Ranges and Tail Risk

Each InfraRel run presently produces a single set of reliability and consequence indices for one defined system configuration and set of input parameters. These point estimates are useful for comparing alternatives, but they do not by themselves describe the complete range of possible outcomes or uncertainty in the assumptions.

Range-based assessments can be developed through a structured matrix of InfraRel cases in which selected failure rates, restoration times, dependency assumptions, demand levels, recovery times and cost parameters are varied. Because individual cases can generally be evaluated independently, parallel processing can make large sensitivity and scenario studies practical.

Sensitivity studies

Vary one parameter at a time to identify influential assumptions.

Coordinated scenarios

Compare defined normal, adverse and extreme combinations of conditions.

Uncertainty sampling

Evaluate sampled parameter combinations to estimate ranges, percentiles, exceedance probabilities and tail exposure.

Running more cases cannot compensate for incomplete system boundaries, omitted dependencies or poorly supported input assumptions. Results should be interpreted together with their assumptions and data limitations.

Use cases

Potential Applications

The following are potential applications where the methodology is being developed and demonstrated. Validation and case studies are still in progress.

Comparing alternative data-center power architectures

Evaluating grid–facility–computational-load dependencies

Identifying single points of failure

Assessing cooling and communications dependencies

Comparing redundancy and restoration strategies

Evaluating UPS, BESS and local-generation roles

Estimating interruption and outage costs

Testing configuration changes and resilience investments

Supporting utility and large-load reliability discussions

Exploring other interconnected critical infrastructures

Limitations & development

Transparent About What the Model Can—and Cannot—Answer

InfraRel can organize complex reliability questions, compare defined alternatives and expose the consequences of failures and assumptions. It does not eliminate uncertainty or provide universally applicable answers. Data-center architectures, workloads, controls, recovery processes and economic consequences differ considerably among facilities.

Limitation Practical response
Uncertain failure and repair data Sensitivity ranges and improved field data
Single-point result per run Structured matrices of parallel cases
Difficult-to-quantify common-mode events Explicit scenarios and bounded assumptions
Changing load and workload behavior Multiple operating and recovery states
Simplified supporting systems Selective detail with explicit dependencies
Facility-specific outage costs Cost ranges and scenario analysis
Limited rare-event evidence Stress cases and tail-risk reporting

The objective is not to claim precision beyond the available evidence. It is to make assumptions explicit, identify which uncertainties influence the decision and test whether conclusions remain valid across credible conditions.

Collaboration

Case Studies and Technical Paper in Development

GR is developing comparative case studies to demonstrate the methodology, required data, reliability indices and configuration tradeoffs. Detailed assumptions, results and limitations will be presented in a forthcoming technical paper.

We welcome opportunities to explore pilot applications and collaborate with utilities, data-center operators, equipment manufacturers, engineering organizations, software developers and researchers who can contribute system knowledge, operational data and validation experience.

Data-center power & grid reliability → · From Analysis to Decision → · Utilities: grid–data center reliability →