Data-center cooling redundancy, N+1 versus 2N: what each topology survives, failure by failure

Every page on data-center redundancy defines N, adds one for N+1, doubles it for 2N and maps the result to a tier. None says which failures each label survives, and on a cooling plant the answer turns on a header, a controller and a water supply the label never counts.

Design-partner scope17 min read

By Jeel Patel, Founder at HVAC Software

The plant drawing in the corridor says N+1. At 08:30 a technician signs a permit and isolates CRAH-07 for its quarterly service, and at 10:40 CRAH-03 trips on a compressor fault. The drawing still says N+1, the BMS graphic still shows nine green units, and the hall is now one unit short of its design load.

N+1 cooling redundancy installs one capacity component, a chiller, pump, tower or CRAH, beyond the N the design load needs, so the plant survives one unit failing or one unit out for maintenance. 2N installs a second complete system that can carry the full load alone, so it also survives the loss of a whole plant or distribution path. Neither label counts the water supply, the plant controller or the header the units share, and the one published availability model of a water-cooled plant puts N+1 and 2N on a shared water supply at the same 99.883 percent.

  • N+1 is a count of units inside one system. 2N is a second system. Only the second version protects the distribution path, and only if the two paths are physically independent.
  • Cheung and Wang's model finds that one redundant chiller lifts a water-cooled plant from 99.679 to 99.883 percent availability, and that a second plant on the same water supply adds nothing measurable, because both plants fail when the water does.
  • Uptime Institute says the tiers do not prescribe N+1 or 2N, that Tier IV is possible with N+1 components, and that the downtime percentages every explainer quotes were never part of the Tier definitions.
  • Redundancy is a live state, not a label. The spare is gone the moment a permit is signed, and Google's 2026 report shows a plant losing its redundant source to construction work and its running plant to a controller that dropped offline.

What is N+1 and 2N cooling redundancy?

N is the number of components that carry the full design load, and Uptime Institute defines it as the number of components that are minimally required to meet the load demand. Everything else is arithmetic on N, and the arithmetic is done per system: a plant can have N+1 chillers, 2N pumps and N+1 CRAH units at the same time.

The labels describe how many units exist, not how they are connected. That distinction is the whole page, because a cooling failure travels through pipes, controllers and a water supply that the unit count never sees.

N
The capacity needed to carry the full design load at the design ambient, with no spare. Nine 100 kW CRAH units for a 900 kW hall.
N+1
One redundant capacity component in the same system, sharing the same header, controls and source. Survives one unit failing or one unit out for maintenance, not both at once.
N+2
Two redundant components in the same system. Cheung and Wang find it adds 0.0001 percent availability over N+1.
2N
A second complete system able to carry the full load on its own. Survives the loss of a whole system, and of a distribution path if the two paths are independent.
2(N+1)
Two complete systems, each with its own spare. Cheung and Wang find it adds 0.0003 percent over 2N.
Distributed redundant
Three or more independent cooling systems, each chiller feeding two or three hall units, so the load redistributes when any one system is lost. Uptime Intelligence records it in the Middle East and Russia.
Concurrently Maintainable
Uptime's Tier III outcome: any capacity or distribution component can be removed from service on a planned basis without affecting the IT load.
Fault Tolerant
Uptime's Tier IV outcome: any single unplanned failure of equipment or a distribution path is stopped short of the IT load.
Continuous Cooling
Stable cooling through the mechanical restart after a utility loss, including the transfer to generators. Uptime requires it for Tier IV only.

Vertiv's position is that N+1 redundancy has become the industry-standard minimum for cooling system design, applied across air handlers, liquid loops, chillers, pumps and controls. Uptime Intelligence records the same default for liquid cooling: CDUs are commonly deployed in an N+1 configuration when redundancy is expected, rather than 2N.

How does a cooling failure travel through an N+1 plant and a 2N plant?

A failure starts somewhere in a chain and moves toward the IT inlet. The chain is the same on every plant: a cooling source, the controls that sequence the plant, the chillers, pumps and towers, the headers and loops that distribute the water, the CRAH or CDU units in the hall, and the inlet temperature the servers see.

N+1 puts a spare inside two boxes of that chain, the plant and the hall units. 2N duplicates the run from controls to hall units, if the drawing is honest about independence. Neither adds a second source unless somebody specified one, which is where the availability model below turns.

The source sits outside every frame. Cheung and Wang's model holds every water-cooled configuration at the water utility's own availability until the second plant gets a different source.

Two of Google Cloud's own incident reports show the chain failing at points the labels do not count. In July 2026 a 3 millisecond voltage drop in the Netherlands meant the chiller controller dropped offline during the voltage transient event, failing to signal the chilled water distribution pumps to restart, the redundant source was unavailable to construction work, and the hall reached 44 °C. In July 2022 in London, a data center could not maintain a safe operating temperature due to a simultaneous failure of multiple, redundant cooling systems combined with the extraordinarily high outside temperatures, and the cooling failure lasted 4 hours 8 minutes.

Neither report names the plant's redundancy label, and this page does not assign one. What both show is that the chillers were not the thing that failed. The controller, the second source and the outside air were.

Which failures does each redundancy topology survive?

This is the table the result set does not have. Each row is one failure, each column is one topology, and the last column says where the verdict comes from, because a verdict without a source is an opinion. Rows marked by arithmetic are unit counting, shown in the worked example further down.

FailureN+1, one plant2N, two plants on one water supply2N, second plant on its own sourceDistributed redundantWhere the verdict comes from
One chiller, pump, tower or CRAH failsSurvives, now at NSurvivesSurvivesSurvives, load redistributesUptime Tier II definition · Cheung and Wang Table 9
One unit isolated for planned maintenanceSurvives, spare consumedSurvives, one system short a unitSurvivesSurvives, but each piping system is not concurrently maintainableUptime Tier II and III definitions · Uptime Intelligence
Maintenance on one unit, then a second unit failsBelow NSurvives, 2N minus two units remainSurvivesSurvives if the remaining systems cover the loadBy arithmetic · Uptime: maintenance raises the risk of disruption
Header, loop or isolation valve failsOutage, unless a second path existsSurvives if the two paths are physically independentSurvivesLoad redistributes to the other systemsUptime Tier III definition · Cheung: headers add 0.00008 percent
Plant controller drops offlineOutage, unless controls are redundant or fail safeBoth plants lost if one controller sequences bothSameEach system has its own controlsStein and Gill · Google europe-west4-a, 2026
Controls lose power on a utility transientPumps do not restart and the chillers stopSame, unless the controls are on UPSSameSameGoogle europe-west4-a: chillers kept power, pumps did not restart
Water utility or makeup water lostOutage once storage runs outOutage: both fail if the water supply becomes unavailableSurvives on the other sourceDepends on each system's sourceCheung and Wang · Uptime: 12 hours of makeup water at Tier III
Outside air above the design ambientEvery unit derates at onceSameSameSameGoogle europe-west2-a, 2022 · ASHRAE A1 envelope
Second plant out for construction or a tie-inNot applicableBecomes N until the work endsBecomes NLoses one system for the durationGoogle europe-west4-a: redundant source unavailable to construction
Utility loss and the mechanical restart gapGap until restart, unless Continuous CoolingSameSameSameUptime: Continuous Cooling required at Tier IV only
Ten failures, four topologies. The first two rows are what the labels promise. The other eight are what the drawing has to show separately.

Read down the N+1 column and the promise is exact: one unit, or one permit, and nothing else. Read down the shared-supply 2N column and half the rows say the same as N+1, because the second plant is bolted to the same water, the same controls power and the same weather. The column that changes the answer is the one with its own source, and the peer-reviewed literature is explicit about why.

On controls, Stein and Gill, writing in the ASHRAE Journal, say the design community agrees on N+1 pumps, chillers and cooling units but that there is far less agreement on and understanding of the redundancy requirements and options for the controls components, and that in their experience more major cooling failures come from poor controls design than from mechanical equipment. Their rule is that the loss of a single controller should not result in the loss of data-center cooling, whatever the mechanical label says.

What does the availability model say about N+1 versus 2N?

Cheung and Wang built a series and parallel model of a four-chiller water-cooled plant from field failure rates and ran it through every common configuration. Their abstract is blunt: it is crucial to install a redundant chiller or redundant chiller plant with alternative cooling sources, while distribution headers and the 2(N+1) configuration do not improve the reliability and availability of data center cooling systems effectively.

ConfigurationChillersReliability after one yearAvailability
N, no spare44.43%99.679%
N+1543.63%99.883%
N+2643.78%99.883%
N+1 with every distribution header543.75%99.883%
2N, second plant on the same water supply844.22%99.883%
2(N+1), same water supply1044.62%99.883%
2N, one plant with a redundant controller845.53%99.884%
2N, second plant on a separate water source898.01%99.9994%
2N, second plant air-cooled899.12%99.9996%
Water utility alone, the input every row sharesnone44.64%99.883%
Cheung and Wang, Tables 7 to 12, quoted to the precision shown. Every water-cooled row on one supply sits at the water utility's own availability, 99.8834 percent.

The table has one story in it. Adding the first spare chiller moves availability from 99.679 to 99.883 percent, and after that nothing on the same water supply moves it again, not a second spare, not a full set of headers, not a second plant of four more chillers. The paper states the mechanism: the plants share the same water supply if both are water-cooled, and both fail if the water supply becomes unavailable.

The rows that move are the ones that change the source. A second plant on separate water, or on air-cooled chillers, takes one-year reliability from 44 to 98 or 99 percent, which the authors describe as a different source of cooling supply doubling the reliability of a cooling system after a year. The unit count in those rows is the same eight chillers as the shared-supply 2N row above them.

Do Tier III and Tier IV require N+1 or 2N cooling?

No, and the standard's owner says so in writing. Every page one result for this query maps Tier III to N+1 and Tier IV to 2N or 2N+1, and Uptime Institute's own myths article says that increasing the component count does not determine or guarantee achievement of any specific Tier level, because the tiers also evaluate distribution pathways, and that it is possible to achieve Tier IV with just N+1 components depending on how they connect to redundant paths.

The tiers are outcomes. Uptime's explainer states that the classification does not prescribe any specific technology, schematic or other design criteria beyond those outcomes, and TIA-942's ratings are written the same way, as capacity components and distribution paths rather than as N+1 or 2N.

Uptime Tier · TIA-942 RatingCapacity componentsDistribution pathsWhat that means for cooling
Tier I · Rated-1Non-redundantSingleDedicated cooling, and a shutdown for preventive maintenance and repairs
Tier II · Rated-2Redundant: chillers, cooling units, pumps, heat rejection equipmentSingleA unit can be removed for maintenance. An unexpected shutdown of the path still affects the load
Tier III · Rated-3RedundantRedundant, with one path required to serve the load at a timeNo shutdown for maintenance or replacement. Twelve hours of concurrently maintainable makeup water. N+1 chillers on A and B loops with one loop normally disabled is permitted
Tier IV · Rated-4Redundant, in physically isolated systemsMultiple, independent, activeAny single failure is stopped short of the load. Continuous Cooling through the mechanical restart. Active/active for all systems, and no single controller or weather sensor shared by the units
Uptime's tiers page and journal, and TIA's own brochure for the ratings. The 99.671, 99.741, 99.982 and 99.995 percent figures every explainer attaches to these rows were removed from the Tier Standard in 2009 and, per Uptime, were never part of the definitions.

Two lines from the standard's owner matter more to a cooling plant than the whole N+1 and 2N vocabulary. The first is that Tier III has no active/active requirements for mechanical systems, so N+1 chillers feeding separate A and B loops with one loop normally disabled is compliant, which is a component count of N+1 delivering a path count of two. The second is that Tier IV requires continuous cooling to make the environment stable, and that Uptime recommends Continuous Cooling at densities beyond 4 kW per rack, regardless of Tier.

Continuous Cooling is a requirement on the restart gap, not on the unit count. Uptime describes Tier IV sites meeting it with very large thermal storage tanks on the chilled water, and its requirement that the thermal environment stay stable for any 15-minute period under the ASHRAE guidelines ties the whole exercise to the envelope: 20 °C in an hour and no more than 5 °C in any 15-minute period at the inlet, with class A1 allowable 15 to 32 °C. The plant has to hold that while a spare starts, however many spares there are.

N+1 · 2N
The label is a count of units on a drawing. The state is what is left after the permit is signed, the fault has tripped and the second plant is behind a construction hoarding.

How does maintenance consume redundancy?

A redundancy label is a design-day statement, and the plant does not spend its life on the design day. It spends it with a unit on a permit, a unit derated by a dirty condenser, a loop drained for a valve change or a second plant isolated for a tie-in. Uptime's own tiers page says it plainly: if the redundant components or distribution paths are shut down for maintenance, the environment may experience a higher risk of disruption if a failure occurs.

The worked example below is a composite, built to show the arithmetic rather than any site. One 900 kW hall, CRAH units rated 100 kW at the design supply temperature, so N is nine. The N+1 hall has ten units on one plant and one loop. The 2N hall has two systems of nine, each on its own plant and loop.

TimeEventN+1 hall: units able to carry load, against N = 92N hall: systems able to carry the full load alone
Week beforePlant B isolated for a construction tie-inNot applicableOne, plant A, for the whole week. The hall graphic shows no fault
08:00All units healthyTen. One spareTwo. One spare system
08:30CRAH-07 isolated on a PM permitNine. Exactly N, no spareSystem A at eight of nine, system B whole. No spare system
10:40CRAH-03 trips on a compressor faultEight. Below N by one unitSixteen of eighteen units running, no system whole, load still covered
10:40Live load that morning is 720 kWEight units carry 800 kW. The live load is covered, the design load is notCovered with nine units to spare
Composite. The permit at 08:30 is the moment the spare goes, on both drawings, and nothing on either drawing changes.

The arithmetic makes two points. Under both labels the spare is consumed by a signature, not by a failure, and the drawing does not know. Under 2N the second event leaves sixteen units for a nine-unit load, while under N+1 it leaves eight for nine, which is the whole difference the extra nine units buy on that morning.

The 720 kW row is the one operators actually live in. A hall below its design N is very often still above its live load, which is why the second failure is survived far more often than the arithmetic suggests, and why it is survived by luck rather than design. That margin belongs to a separate calculation on capacity headroom, and this page does not compute it. The point here is that the preventive maintenance permit is a redundancy decision, and the checklist item that needs the unit off has to say what covers it.

Google's 2026 report is the same ledger written by an operator. The redundant cooling source was not available due to known ongoing construction work at the facility, so the site was running on one plant by decision before the utility transient arrived, and the transient took that plant's pumps through the controller. The remediation list includes an interim portable UPS for the chiller controllers and an investigation of chiller pump control system redundancy, neither of which is a chiller count.

What do the BMS, DCIM and CMMS already hold about redundancy state?

Each system holds one input to the state and none holds the state. The gap is not that the systems fail to talk. It is that installed count, isolation, faults and derating live in four places, and the subtraction is done by a person on a phone, if it is done at all.

SystemWhat it holds about redundancy stateWhat it does not hold
BMS or BASUnit run status, alarms, valve positions, the restart sequence and the flow readings Uptime says are needed to prove N+1Which unit is under a permit or a lockout. A derated unit still reads running
DCIMThe asset list, the design label, the capacity per unit and the hall's environmental historyLive isolation. It knows the plant is N+1, not that N+1 ended at 08:30
CMMS or EAMThe PM permit, the lockout record and the open work order on the faulted unitCapacity. It cannot say whether the units left are enough for N, or for the live load
The commissioning recordThe proven N: which units actually carried the load in Level 5 testing with a chiller removed and a loop isolatedAnything after handover, unless the readings were kept as values rather than a report
The plant drawingThe labelThe state at 10:40
Five holders, four inputs, no subtraction. The state is installed minus isolated minus faulted minus derated, against N, and nobody owns the minus signs.

The commissioning record deserves its own line because it is the only place N was ever proven rather than drawn. Level 5 integrated testing removes redundant chillers, pumps and CRAH units under load to show what remains still holds the hall, and the commissioning checklist is where those readings should have been kept as values. A plant whose N has never been demonstrated with a unit removed has a label, not a redundancy.

How to write your live redundancy state down this week

The formula is subtraction. The work is deciding what goes into each term, and writing it where the shift can see it. The order below produces a state per hall and per plant that can be compared with the drawing.

  • Write N per hall and per plant from the commissioning record, not from the drawing. If Level 5 testing never removed a chiller under load, N is a design figure and the state carries that caveat.
  • For every spare, write what it is a spare for: a unit only, or the path too. Trace the header, the loop and the isolation valves to find out whether a second path exists or whether the drawing only doubled the units.
  • Name the shared dependencies the label does not count: the water utility and makeup storage, the controls power supply, the plant controller, the weather sensor and the outside air. Each one is a row in the failure table with its own verdict.
  • Compute the state each shift as installed minus isolated minus faulted minus derated, against N, and write the number on the shift log beside the label. A derated unit counts at its measured capacity, not its nameplate.
  • Make every permit a redundancy decision. A permit that takes a unit off writes the new state, names the unit that covers it, and says whether a second event is now an outage.
  • Define the moment the spare is consumed as a condition with a consequence, in the alarm philosophy, so that a permit plus a trip raises something a person has to answer rather than two green rows going amber.
  • Prove the state at least once a year with the redundant unit removed under load, the way Level 5 testing did at handover, and keep the readings as values.

The state is also the input the failure-time ranking assumes. How long a failure gives you depends on what is left running when it happens, and what is left running is the number this list produces.

Who is allowed to consume redundancy?

A named person does, by signing the permit, and the signature should carry the state it creates. The technician who isolates CRAH-07 owns the unit. The shift lead who approves the permit owns the fact that the hall is now at N, and the decision about whether the second event, a trip on any other unit, is now tolerable or is now an outage.

CW

Coolant flow restriction upstream of CDU-2E

CW-2179CriticalAwaiting approvalDecided 03:15:44
Observed
CDU-2E Δp 0.8 → 0.3 bar over 90 s · RACK-14/16 inlet +6.2 K · hall return +2.1 K
Inferred
Coolant flow restriction upstream of CDU-2E. Rack symptoms are downstream, not independent faults.
Confidence
Medium — consistent with two prior events on this loop; no flow meter on the affected branch.
At risk
Loop 2E · 14 racks · N+1 already consumed by scheduled work on CDU-2F.
Margin
≈ 11 minutes of thermal headroom at current load before inlet exceeds class limit.
Response
Qualified on-site engineer with CDU authorisation · MOP-114 attached · escort not required.
Approval
Requires named human approval before dispatch. The system does not act on this alone.
Closes when
Δp holds ≥ 0.7 bar for 30 min AND rack inlet returns to baseline AND the engineer confirms the cause found.
The state, drawn: the permit, the trip, the units left against N, and the moment the spare went, on one record instead of four.

What the layer would do is read. It would hold the BMS unit states and the CMMS permits against N, mark the 08:30 permit as the moment the spare went, and put that state on the 10:40 work order so the dispatcher sees a hall below N rather than a compressor alarm. It would not start, stage or restart a chiller, pump or CRAH, open a valve, or change a permit, and it would not decide whether the work proceeds.

What this cannot do

A page assembled from Uptime's own articles, TIA's brochure, one reliability paper and two operator incident reports is not your plant's failure analysis. The failure table's verdicts are the sources' verdicts, and the rows marked by arithmetic are unit counting on a composite hall. Cheung and Wang's numbers belong to their four-chiller model, and the Tier Standard itself was not opened for this page, only the standard owner's published summaries of it.

Nothing here predicts a failure or diagnoses one. A decision layer that holds the state can say that hall B has been below N+1 since 08:30 and below N since 10:40, and it can say which shared dependency the drawing never counted. It cannot say which unit fails next, and it does not start, stop, stage or reset anything. A named person signs the permit, and a named person decides whether the work goes ahead.

Answered

Which is better for data-center cooling, N+1 or 2N?

2N survives more failures, and only some of them are the ones you will have. It covers the loss of a whole plant or distribution path, which N+1 does not. It does not cover a shared water supply, a shared controller, controls power or outside air, and Cheung and Wang's model puts 2N on one water supply at the same 99.883 percent availability as N+1. The better answer is the topology whose drawing shows a second source and independent paths.

Is N+1 enough for data-center cooling?

N+1 is the industry-standard minimum for cooling, in Vertiv's words, and it covers one unit failure or one unit on a permit. It is enough while the spare is intact and the failure is a unit. It is not enough on the morning a permit is signed and a second unit trips, and it says nothing about the header, the controller or the water supply. Whether that is acceptable is the risk decision Uptime says each owner has to make.

Does Tier III mean N+1 and Tier IV mean 2N?

No. Uptime Institute's own guidance says component count does not determine tier level, that Tier IV can be achieved with N+1 components on properly redundant distribution paths, and that Tier III has no active/active requirement for mechanical systems. The tiers are outcomes, Concurrent Maintainability and Fault Tolerance, measured on capacity components and distribution paths together. The downtime percentages attached to them in most explainers were never part of the definitions.

What is 2N+1, and is it worth it for cooling?

2N+1 is two complete systems plus one further spare, and 2(N+1) is two complete systems each with its own spare. Cheung and Wang's model finds 2(N+1) adds 0.0003 percent availability over 2N on the same water supply, which the authors call not necessary. The money buys far more as a second cooling source, which in the same model takes one-year reliability from 44 percent to 98 or 99 percent.

What is distributed redundant cooling?

Distributed redundant cooling is three or more independent cooling systems, each chiller feeding two or three hall units, so that any one system can be lost and the load redistributes. Uptime Intelligence records it mainly in the Middle East and Russia, notes that it avoids single points of failure on the chilled water side, and notes the cost: every chiller stays on at all times to keep the redundancy, and each piping system is not itself concurrently maintainable.

Can a 2N cooling plant still have a single point of failure?

Yes, and the published cases are the shared ones. A single controller sequencing both plants, controls that lose power on a utility transient, a water utility both plants draw from, and outside air both plants reject heat into are all single points a 2N label does not count. Google's 2026 report shows a controller dropping offline and stopping the pumps while the chillers kept power, with the redundant source already out for construction.

What is the difference between redundancy and concurrent maintainability?

Redundancy is a count of spare capacity components. Concurrent Maintainability is Uptime's Tier III outcome: any capacity or distribution component can be removed from service on a planned basis without affecting the IT load, which needs isolation valves, a second path and a procedure as well as a spare unit. A plant can be N+1 and still need a shutdown to change a header valve, and Uptime says ductwork and piping count.

Data centerRun data-center cooling as an operation, not an alarm feed.