Data center cooling systems, ranked by how long they give you after a failure
Every guide to data center cooling systems ranks them by efficiency. The number that matters at 03:14 is how many minutes each one leaves you when it stops, and that order is set by where the cold sits, not by the PUE.
By Jeel Patel — Co-founder, FieldCamp · CEO, Upper Route Planner
Utility power drops at 03:14. The UPS holds every server in the hall, and it does not hold the fans, the pumps or the chiller, because in most facilities those sit on the generator. For the sixty seconds until the generator picks up, every kilowatt of IT load is heating the air in the room with nothing moving it.
Data center cooling systems are the air and liquid architectures that carry heat from the chip to the outdoors, and they divide into CRAC units with their own refrigeration, CRAH units on a chilled-water loop, rear-door heat exchangers, direct-to-chip cold plates and immersion tanks. Ranked by the time they leave you after their own worst failure, the order runs from a cold-plate loop that loses flow, where a processor can throttle in 23 seconds, to a chilled-water plant with stored water, which can ride through a 10 to 15 minute chiller restart. What sets the order is where the cold thermal mass sits relative to the failure.
- The efficiency ranking every guide uses and the failure-time ranking are different lists, and for chilled water they run in opposite directions.
- Without airflow the room can heat at 5 °C a minute or more, and ASHRAE allows no more than 5 °C of change in any 15 minutes.
- A chilled-water plant's time is its chiller restart, 10 to 15 minutes standard, and stored water is what bridges it.
- A cold-plate loop's time is its pump: 23 seconds to throttling at full load without active pump redundancy.
What are the data center cooling systems, in one table?
Schneider's taxonomy counts 13 fundamental heat removal methods between the IT space and the outdoors, and they collapse into six architectures a facilities engineer actually runs. Each one is defined by two things: where the refrigeration cycle sits, and what carries the heat from the hall to it.
The third column is the one no comparison page carries, and it is the one the rest of this guide is built on. It says what is cold and near the load at the moment the architecture stops, because that is what buys the time.
| Architecture | Where the cold is made | What carries the heat out | Cold mass near the load when it stops |
|---|---|---|---|
| DX CRAC | In the unit, by its own compressor | Refrigerant to a condenser or glycol to a dry cooler | Room air and the coil, seconds' worth |
| Chilled-water CRAH | In a chiller plant, at about 8 to 15 °C | Chilled water in pipes to the CRAH coil | Water in the pipes, minutes' worth, plus any storage tank |
| In-row and containment | As CRAC or CRAH, placed in the row | As its parent architecture | Less room air than a perimeter design, so faster either way |
| Rear-door heat exchanger | Chiller or cooling tower via a plate exchanger | Water through a coil in the rack door, 4 to 15 gallons a minute | The room's own cooling, once the door goes warm |
| Direct-to-chip cold plate | Facility water, through a CDU | Coolant in a technology loop to the plate | Almost none: the loop volume, seconds' worth |
| Immersion | Chiller or tower via a heat exchanger | Dielectric fluid in the tank | The whole tank of fluid |
Economizers are not a seventh architecture but a mode of the chilled-water and CRAC ones, and they matter here for one reason. When the refrigeration fails and the weather allows, an economizer carries the load without it, and Uptime has certified Tier III sites that supply 27 °C air with no refrigeration or DX at all.
Why does failure time decide the ranking, and not efficiency?
Because the four things done to a hall in the name of efficiency all shorten the window. Schneider's White Paper 179 names them: right-sizing the cooling to the load, raising rack density, raising supply and chilled-water setpoints, and containing the aisles. Each is correct under utility power and each removes thermal mass or reserve capacity from the moment the power fails.
The rate is the problem. With no airflow at all, the paper puts the rate of rise at 5 °C a minute or more depending on density and layout, while the ASHRAE thermal guidelines allow no more than 5 °C of change in any 15-minute period and 20 °C in an hour. The first minute of a power loss with the fans on generator violates the guideline on its own.
The efficiency ranking asks how much power the cooling uses when it works. The failure-time ranking asks how much cold is left in the room when it does not.
Intel's own calculation for a 310 watts per square foot hall, published in 2007, was 18 seconds from 72 °F to 105 °F and 35 seconds to 135 °F once airflow stopped. Densities have gone up since, not down, and a 2026 Cundall model of a high-density hall without mechanical UPS or buffer vessels breached the ASHRAE thresholds within seconds.
Which data center cooling system fails fastest?
The table ranks each architecture by its own binding failure, fastest first, with the sourced time and the one thing that buys more. Times are from the sources named in each row and are stated the way the source stated them: an experiment, a model of a named hall, or an operator's own event.
| Rank | Architecture and the failure that ends it | Time you have | What buys more | Source |
|---|---|---|---|---|
| 1 | Direct-to-chip cold plate, total loss of technology-loop flow (CDU pumps stop) | 23 s to processor throttling at 100% utilisation, 56 s at 25%, 305 s idle | Active N+1 pump redundancy in the CDU, UPS on the loop pumps, load migration inside the time-to-throttle | Experiment, Alkharabsheh et al.; cited by ASHRAE TC 9.9 |
| 2 | Any air-cooled hall, loss of airflow (fans on generator, not UPS) | Seconds at high density: 18 s to 105 °F at 310 W/sq ft; at least 5 °C a minute in the general case | CRAH or CRAC fans on UPS for the 60 s to generator | Intel calculation; WP 179 model |
| 3 | DX CRAC, power loss | Fans back on generator at 60 s, then the compressor waits its short-cycle timer, three minutes on one widely installed Liebert model, and the outdoor heat rejection must be on backup power too | Fans on UPS; generator on the whole system including dry coolers; a chilled-water coil in the unit | WP 179; Liebert manual |
| 4 | Rear-door heat exchanger, loss of water flow | The room's time, not the door's: the rack reverts to an air-cooled rack and the room CRAC or CRAH compensates as far as it can | Room cooling sized to carry the rack without the door | Follows from the DOE FEMP case study |
| 5 | Chilled-water CRAH, power loss, fans and pumps on generator, no storage | Pipe water carries the coil for minutes; the chiller restart of 10 to 15 minutes is the binding wait. Modelled hall: over the allowable for about 3 minutes, high until 17 minutes | Fans and pumps on UPS removes the first spike; a quick-start chiller cuts the restart to 4 to 5 minutes | WP 179 Example 1 |
| 6 | Chilled-water CRAH with thermal storage, fans and pumps on UPS | Through the chiller restart, staying under the 32 °C allowable in the model when the tank is sized for the density; in Intel's 2006 outage, through a multi-hour loss with staff arriving 45 minutes later to a normal room | Tank sized for the density and the restart time | WP 179 Figure 4; Intel IT |
| 7 | Immersion, loss of pump or heat rejection | Set by the heat capacity of the fluid in the tank. No operator-published figure was found, so no number is given here | Fluid volume; pump redundancy | Mechanism only |
Read rows 5 and 6 against the efficiency ranking and the reversal is plain. Chilled water is the architecture the comparison pages rank most efficient at scale, and it is also the one with the most time after a failure, because its cold mass is water in pipes rather than air in a room, and water in pipes can be added to.
Read rows 1 and 2 together and the second finding appears. The two fastest failures are not the two least efficient systems but the two with the least cold near the load: a cold-plate loop with almost no volume, and any air-cooled hall in the minute before its fans return.
How does each cooling system actually lose the hall?
Every architecture has a chain of things that must restart, in order, before cooling returns. The chain, not the equipment, is what sets the time, and for a chilled-water plant it looks like this.
- DX CRAC
- The compressor stops with the power and cannot restart until its short-cycle timer clears and the condenser or dry cooler is back on power, which WP 179 calls the least emergency potential of the air systems. Glycol-cooled units without a free-cooling coil get nothing from pumps on UPS, because the fluid is useless until the compressor runs.
- Chilled-water CRAH
- The chiller's controls ride through only a quarter of a cycle, about 4 ms at 60 Hz, so any real dip forces a restart. Intel found sags shorter than a second and voltages below about 85% of nominal enough to trip its centrifugal plant, which then took three to six minutes to come back.
- Contained aisles
- With row coolers off UPS and the doors shut, exhaust recirculates to the inlets through every leakage path and the rise is faster. With perimeter coolers the cold aisle is a cold mass and containment helps, so the same containment can shorten or lengthen the window depending on where the coolers are.
- Rear-door heat exchanger
- The passive door has no moving parts and the server fans push exhaust through it, so when the water stops the fans keep running and the rack simply discharges warm air into the room. The door needs water flow and the room's central cooling is what compensates.
- Direct-to-chip
- The technology loop holds almost no cold volume, so loss of flow reaches the chip in seconds. ASHRAE's bulletin reports a 27 °C jump in loop temperature from a CDU pump changeover made without active redundancy, and the experiment it cites found that losing one pump of a pair, leaving about 68% of flow, had no effect at all.
- Immersion
- The tank is the thermal mass, and it is the largest cold mass any architecture puts next to the load. What it is worth in minutes depends on fluid volume and load, and no operator has published the number.
The DX and chilled-water rows explain why the ISA-18.2 guide's fourth alarm attribute, time to respond, cannot be a fixed number written into a template. It is the remaining length of a restart chain, and the chain is different for every architecture and every failure.
How much stored water does the chiller restart need?
For chilled water the answer has been modelled. WP 179's Table 2 sizes the tank that keeps a hall under ASHRAE's 32 °C A1 allowable through the restart, for four densities and two chiller types, with fans and pumps on UPS.
| Rack density | IT load | Standard chiller, 12-minute restart | Quick-start chiller, 5-minute restart |
|---|---|---|---|
| 1 kW/rack | 100 kW | 0 gal | 0 gal |
| 4 kW/rack | 400 kW | 1,300 gal (4.9 m³) | 350 gal (1.3 m³) |
| 8 kW/rack | 800 kW | 3,800 gal (14.4 m³) | 1,300 gal (4.9 m³) |
| 12 kW/rack | 1,200 kW | 7,000 gal (26.5 m³) | 2,500 gal (9.5 m³) |
Two things in that table transfer to every architecture. A faster restart is worth more than a bigger tank, and above roughly 4 kW a rack the window without either is short enough that Uptime recommends continuous cooling regardless of Tier, though only Tier IV requires it.
What do the BMS, DCIM and CMMS already know about the window?
Each system in the building holds one input to the failure-time question and none of them holds the answer. That is not a gap in any of them, because none was asked to compute a window.
| System | What it holds that the window needs | What it does not hold |
|---|---|---|
| BMS / BAS | Supply and return temperatures, chiller and CRAC status, alarms, the setpoint | Which loads are on UPS, what the chiller's restart time is, how much water is in the pipes |
| EPMS | Which fans, pumps and chillers are on UPS, generator or utility | What that connectivity is worth in minutes |
| DCIM | Rack density, the hall's volume, which racks a CRAH serves | The restart chain, or the state of the thermal storage |
| CMMS / EAM | OEM literature, chiller and CRAC service history, the restart figure if anyone typed it in | Live temperatures, or the failure that is happening |
The join is exactly what FM Global's data sheet asks a facility to make by hand, and it is the reason the number is usually in nobody's head at 03:14. It is also the reason the interval between the alarm and the dispatch is spent guessing how urgent the alarm is.
How to find out which failure ends your hall first
The procedure below is FM Global Data Sheet 5-32's loss-of-cooling contingency plan, as reproduced by ASHRAE TC 9.9, reordered to answer the ranking question for one hall. It produces a document, not a number, and the number is a separate exercise.
- List every cooling component whose breakdown stops cooling: chillers, chilled-water and condenser-water pumps, tower fans, air-handler fans, control valves that fail closed, controls, variable-speed drives, and the breakers feeding each. This is the data sheet's own list.
- For each, write down what powers it during a utility loss: UPS, generator, or nothing until the generator. The EPMS holds this, and it is the single fact that most changes the ranking.
- Get the restart time from the OEM literature for every chiller and CRAC, including any short-cycle or anti-recycle timer. Put it in the CMMS against the asset, not in a binder.
- For liquid-cooled racks, record whether the CDU has active N+1 pump redundancy and whether the technology loop pumps are on UPS, which ASHRAE's 2024 bulletin makes its second and third design recommendations.
- Have a qualified engineer calculate the rate of rise for the hall at its current density, as the data sheet requires, and the time to the OEM's damage threshold for the most sensitive equipment in it.
- Run the three scenarios the data sheet names, a one-second, one-minute and one-hour utility loss, plus the single-component breakdowns, and record for each which restart chain is binding.
- Set the BMS to alarm on rate of change as well as on level, no more than 2 °C above the setpoint, and give the controls their own battery backup so the alarm survives the event it reports.
The row that comes out binding is your hall's rank in the table above, and it will not always be the architecture you expected. A chilled-water hall whose pumps sit on generator only is a row 5 hall, and adding a tank without moving the pumps to UPS gains nothing in the first minute, because the tank cannot be used until the pumps run.
Who decides to shut load down, and when?
A person, named in advance. The data sheet's implementation clause says to designate at least one person per shift with the authority to implement the plan, including the power-isolation plan, when shutting equipment down is what prevents damage, and to practise the decision path annually.
That clause is the human boundary this whole site is built around, and the ranking is the input that person needs. Knowing the hall is a row 2 hall changes what the on-call engineer does in the first minute, and knowing it is a row 6 hall changes whether anyone is woken at all.
What this cannot do
A ranking is not a ride-through figure for your hall. The times in the table are an experiment on one rack, a model of one 400 kW room, and one operator's 2006 outage, and each is labelled as such in its row. Your number depends on your density, your pipe volume, your restart chain and which loads are on UPS, and only the calculation the data sheet asks for produces it.
Nothing here takes control authority over any of it. A decision layer that reads the architecture and the restart chain can say how long the hall has and who is authorised to act, and it does not start a chiller, open a tank valve, or move a setpoint. The person with the authority per shift does that, and the layer records what they decided and whether the telemetry agreed.
Answered
What are the main types of data center cooling systems?
Data center cooling systems are air-cooled or liquid-cooled. Air cooling uses CRAC units with their own refrigeration or CRAH units fed by a chiller plant, usually with raised floors and contained aisles. Liquid cooling uses rear-door heat exchangers on the rack, direct-to-chip cold plates fed through a coolant distribution unit, or immersion in a tank of dielectric fluid. Economizers are a mode of the air systems, not a separate type.
Which data center cooling system is most reliable when it fails?
A chilled-water system with stored water and its fans and pumps on UPS gives the most time after a failure, because water in pipes and tanks is cold mass that can be added to. Schneider's model kept a 400 kW hall inside the acceptable range through a chiller restart with those three measures, and Intel's tanks carried a real multi-hour outage in 2006.
How long does a data center have after cooling fails?
It ranges from seconds to the length of a chiller restart. A cold-plate loop that loses flow can throttle a fully loaded processor in 23 seconds, an air-cooled hall with no airflow can heat at 5 °C a minute or more, and a chilled-water plant with fans and pumps on backup power has until its chiller restarts, typically 10 to 15 minutes. Your figure needs your density and your restart chain.
Why does a chiller take so long to restart after a power loss?
A chiller's controls ride through only a quarter of a cycle, about 4 milliseconds at 60 Hz, so almost any dip forces a restart. On restart the chiller checks its controls, compressor, oil and water systems before loading, which takes 10 to 15 minutes on a standard machine and 4 to 5 on a quick-start design. Intel reported three to six minutes for its centrifugal plant.
Does N+1 cooling redundancy give you more time after a power loss?
Not by itself. N+1 protects against one unit failing while the rest run, but after a utility loss every chiller restarts together and the spare is restarting with them. Time after a power loss comes from what is on UPS, from stored water and from restart speed, which is why Uptime requires continuous cooling only at Tier IV and recommends it above 4 kW a rack regardless of Tier.
What happens to a rear-door heat exchanger if the water stops?
The rack goes back to being an air-cooled rack. A passive door has no moving parts and the server fans push exhaust through it, so with no water the fans keep running and the rack discharges warm air into the room. How long that is survivable is set by the room's central cooling, which the DOE case study notes can compensate for warm door discharge as far as its capacity allows.
How fast does a direct-to-chip cooled server overheat if the pump fails?
Fast enough that the ASHRAE bulletin treats it as its first concern. The cited experiment found a fully loaded processor began throttling 23 seconds after total loss of flow, 56 seconds at 25% utilisation and about five minutes idle, with a 20 °C coolant set point adding under 50%. Losing one pump of a redundant pair, leaving about 68% of flow, had no effect.