Direct-to-chip cooling leak: the first ten minutes, and who is allowed to do what

Every page about direct-to-chip cooling says to install leak detection and be ready. None of them says what the server, the rack and the CDU do on their own when the sensor trips, what a person has to decide, or who that person is allowed to be.

Design-partner scope15 min read

By Jeel Patel, Founder at HVAC Software

It is 02:40 and the tray in slot 5 of a liquid-cooled enclosure on row C has gone dark on its own. The enclosure's management module has logged a leak-sensor event against that tray, the CDU at the end of the row is showing a make-up water alarm, and the on-call engineer is reading both from a phone in the car park.

A direct-to-chip cooling leak is handled in three layers, and the first ten minutes are about not confusing them. The server's own BMC powers the leaking node off within seconds, and on Lenovo's Neptune systems it blocks that node from powering on again until the leak is fixed. Isolating the rack at its manifold valves is a decision a trained person makes, or a solenoid valve makes for them, and it dries one rack. Stopping the CDU is the last resort, because it takes flow from every rack on the loop and a loaded processor can throttle 23 seconds later, and nobody opens a wet tray until the power cords are out.

  • Layer one is automatic: the BMC powers the node off. Layer two is a choice: isolate one rack at its manifold valves. Layer three is the last resort: stop the CDU, which dries every rack on the loop.
  • The OCP severity matrix decides which layer a leak belongs to: minor is no visible spill, major is coolant inside one server, critical is pooling outside the rack.
  • The person who closes a valve or opens a tray is a trained person as defined by IEC 62368-1, and the leaking water loop, manifold or quick connect is a field replaceable unit that the OEM's technician replaces, not a field repair.
  • Vertiv's CoolChip CDU ships with leak detection set to alarm only, with a ten-second delay, and shuts itself down within a second only when the loop has lost water and flow.

What is a direct-to-chip cooling leak, and where does it come from?

Direct-to-chip cooling puts a liquid-filled cold plate on the processor, so heat leaves the chip in coolant rather than air. That coolant runs in a technology cooling system loop that a coolant distribution unit keeps separate from the building's facility water, and the CDU's own heat exchanger is the only place the two loops meet. The coolant in most deployments is water with about 25% propylene glycol and an inhibitor package, which one supplier describes as non-toxic at loop concentration and easy to dye.

A leak is that coolant anywhere it should not be, and the OCP leak-intervention paper's rule is that each connection point is a potential risk: the cold plate, the tray's quick disconnects, the rack manifold, the hoses and the CDU. The same paper records two lessons from US national laboratories. A cheap O-ring dripped for months without ever tripping the facility leak system, and quick disconnects wept because the weight of hanging hoses deformed their seals.

Cold plate
The copper block on the processor. In the Binghamton failure experiment it carried 75% of the chip's heat, which is why losing its flow reaches the chip in seconds.
Rack manifold
The supply and return headers that feed each tray. Its shutoff valves are the isolation point for one rack, and whether they are hand valves or solenoids decides who can close them.
Quick disconnect
The dry-break coupling at the tray. The OCP Universal Quick Disconnect is specified to drip under 0.063 cc when broken at 200 psi, so the coupling is rarely the flood.
Leak detection rope
A two- or four-wire cable whose resistance drops when conductive coolant bridges it. Four-wire ropes report where along the cable the leak is.
Coolant distribution unit
The pumps, heat exchanger, level sensor and leak tape at the end of the row or in the rack. It keeps the loops apart and it is the one thing you do not stop first.

The OCP paper also gives the leak a size, and every step below hangs off that size. Its coolant leak severity matrix is reproduced here with its own column names, because it is the trigger for the sequence.

SeverityCoolant liquid lossDetection tripImpact to IT gear
MinorNo visible spillIndirect detection: a pressure or reservoir-level change at the CDUNone
MajorVisible within the IT equipmentDirect detection in a single zone, usually the tray's own rope or the drip tray under itReduced performance or shutdown of a single server
CriticalPooling outside of the rackDirect detection in multiple zones, or outside the rackReduced performance or shutdown of multiple devices
The OCP coolant leak severity matrix, Table 1 of the Leak Detection and Intervention paper. Minor is watched, major isolates a server, critical isolates a rack.

What happens in the first ten minutes of a cold-plate leak?

The sequence below is the OEM leak workflows and the OCP options list put in the order the clock runs them. Row C, slot 5 is a composite, assembled from the Lenovo N1380 troubleshooting chapter, the Vertiv CoolChip alarm table and design-partner conversations, and it is written to be recognisable rather than to report a site.

02:40
The tray's leak sensor bridges. Its BMC powers the node off and, on this platform, blocks power permission until the leak is resolved. The enclosure module logs the tray-level event and the BMS raises it.
02:41
The CDU's internal leak tape is dry, but its make-up pump has been running and it raises a low-level alarm. Vertiv's default for both of those is alarm only, not shutdown, so the other eleven racks on the loop keep their flow.
02:43
The shift engineer reads the severity off the matrix. One tray, one rope, no pooling outside the rack: major, not critical. The correct isolation point is the rack's own manifold valves, not the CDU.
02:46
A trained person closes the rack's supply and return valves and pulls the power cords from the enclosure's power conversion stations, which Lenovo requires before any tray comes out and for at least two minutes. One rack is now dark and dry.
02:50
DCIM shows the rack was carrying 34 kW of a 110 kW row. The CDU stays up, the load is migrated off the dark rack, and the OEM service call for the water loop is logged as a field replaceable unit, not a field repair.
Closes when
The reported tray and its two neighbours have been pulled and inspected, the leaking loop, manifold or quick connect has been replaced rather than resealed, the enclosure powers on without shutting itself down again, and the CDU's level and pressure alarms have stayed clear.
Two of the five steps are automatic. The other three are a person's, and the person has to be the right one.

Read the sequence and the thing to notice is where the minutes go. The automatic step took seconds and the CDU did nothing, by design, and the six minutes between the alarm and the closed valve were spent by one person deciding which layer this leak belonged to and whether they were allowed to act on it.

Which leak actions happen by themselves, and which need a person?

The OCP paper divides intervention into manual, which is closing valves and disconnecting hoses, and automatic, which is de-energising the IT equipment and shutting off the coolant. Its definition of an automatic reaction also says who acts at each level: the BMC at the node, the rack manager or the rack CDU at the rack, and the row's CDU talking to the building management system at the row.

ActionWho or what takes itTriggerWhat it costsSource
Power the leaking node offThe server's BMC, automaticallyThe tray's own rope or drip sensorOne node. Lenovo's core module also blocks power permission until the leak is fixedLenovo SR780a V3
Cut DC power to the tray or rack before it enters the chassisThe rack manager or power shelf, on a signal from the BMC or the BMSA tray-level leak on a rack-scale systemOne tray or one rack. NVIDIA's fault table lists power-shelf DC off, rack breakers off and liquid valves shut in both directionsNVIDIA Mission Control BMS integration
Isolate one rack at its manifold valvesA trained person at a hand valve, or a solenoid on a BMS signalA major leak on the matrix, or two sensors agreeingOne rack goes dry. A hand valve isolates one side of the network only, and opening an un-terminated valve is a catastrophic leakOCP Leak Detection and Intervention
Pull the power cords from the enclosureA trained personWater seen outside the enclosure, or before any tray is removedThe enclosure. Lenovo requires at least two minutes off to avoid the power stations latchingLenovo N1380 guide
Break the tray's quick disconnectA trained personTray removal for inspectionUnder 0.063 cc on a UQD2 at 200 psi by specification, if the valve is not held open by contaminantOCP paper, Danfoss UQD sheet
Shut the CDU downThe CDU itself if configured to, or a personVertiv: a substantial leak on its tape with the unit set to alarm plus shutdown, or loss of level with flow under half of setpointEvery rack on the loop loses flow. A loaded processor throttles in 23 secondsVertiv CoolChip CDU100kW manual, Alkharabsheh et al.
Assembled from the OCP paper, two Lenovo guides, NVIDIA's BMS integration page and the Vertiv CDU manual. Each row names the equipment it describes.

Two rules from the OCP paper sit under that table. There is an inverse relationship between operational impact and potential damage, so the action that protects the most hardware is the one that takes the most of it offline. And on nuisance alarms the paper's remedy is redundant sensors: if either detects fluid, alarm the operator, and only if both do, intervene automatically.

The BMC acts in seconds and the CDU can act in one. Neither of them knows whether the leak is a drip on one tray or a flood under twelve racks, and that judgement is the whole of the first ten minutes.

Why is stopping the CDU the wrong first move?

Because it is the one action that reaches every rack. The OCP paper calls stopping flow at the CDU the extreme example of a leak response, one that requires every liquid-cooled server on that CDU to be shut down, and notes that a row CDU often feeds ten or more racks and that a redundant CDU must receive the same shutdown signal or it will simply keep pumping.

What a stopped loop does to a processor is measured. In the Binghamton experiment that ASHRAE's resiliency bulletin cites, a fully loaded processor began throttling 23 seconds after total loss of flow, 56 seconds at 25% utilisation and about five minutes idle, which is why the failure-time ranking puts the cold-plate loop first. The OCP paper adds a requirement most sites have never checked: before any CDU shutdown, thermal throttling must be enabled on every cold-plated chip, or the stalled coolant can boil, and a safe shutdown should be issued to the affected servers before the pumps ramp down.

23 seconds
From loss of flow to a throttling processor at full utilisation, in the experiment ASHRAE cites. A CDU shutdown starts that clock for every rack on the loop.

The manuals already lean this way. Vertiv's CoolChip CDU ships with its internal leak tape and the optional under-floor tapes set to alarm only, with a ten-second delay, and the operator has to choose alarm plus shutdown at commissioning. The one leak case where that CDU stops itself within a second is when the level sensor is open and flow has fallen under half of setpoint, which is a loop that has already lost its water.

The same manual's front page tells you where Vertiv would rather the automation sat. It recommends a monitored fluid detection system wired to close field-installed supply and return shutoff valves, which is the rack or row manifold, not the pumps. ASHRAE's bulletin makes the parallel point for pump changeover: without active redundancy on the N+1 pumps, ITE manufacturers have seen a 27 °C jump in loop temperature from the changeover alone, and a loop with no cold volume has no ride-through to absorb it.

Who is allowed to isolate the rack, pull the cords and open the tray?

Lenovo's enclosure guide answers the person question in one sentence: the equipment must be serviced by trained personnel as defined by IEC 62368-1, installed in a restricted access location whose access is controlled by the site's own authority. That standard defines three classes of person, ordinary, instructed and skilled, and a leak puts conductive water next to a power conversion station, which is the energy source the standard has in mind.

The OEM adds its own tier on top of the person class. Lenovo lists the water loop, the manifolds and the leakage sensor as field replaceable units that only trained service technicians install, and its leak procedure says a manifold or quick connect found leaking is discarded and replaced, not resealed. Vertiv's CDU manual restricts installation, service and maintenance to trained and qualified personnel and requires the unit switched off and disconnected before maintenance begins.

  • Equipment expertise on this rack. Someone who can read the enclosure's power-station LEDs, which on Lenovo's N1380 distinguish a power-station leak from an enclosure or tray leak, and who knows which manifold valve is which.
  • OEM authorisation. The loop, the manifold and the sensor are FRUs, so the person who replaces them is the OEM's or is authorised by the OEM, or the warranty is at risk.
  • Headcount and time. Lenovo's tray-leak procedure removes the reported tray and the trays either side of it, and its enclosure-leak test pulls every tray out 10 cm and waits ten minutes before power is reconnected.
  • Site clearance for the row, plus the slip and electrocution hazards the OCP paper names for a significant leak, because the person reaching for the valve is standing in it.
  • A procedure that puts the power cords out before any tray is touched, with the lockout-tagout (LOTO) that governs the loop, which is the same rack-power-off principle the rear-door heat exchanger manuals write down.

Three organisations usually own a piece of this. The in-house engineer closes the valve and pulls the cords, the server OEM's technician replaces the water loop, and the CDU's vendor is the one allowed inside the unit if the leak tape that tripped was its own. None of them is in the other's system, and the leak does not wait for that to be sorted out.

What do the BMC, BMS, CDU controller and DCIM already know about the leak?

Each of them holds one input the ten-minute decision needs, and none of them holds the decision. The OCP paper is explicit that in all cases a leak alarm must reach the DCIM or BMS, and it lists Modbus RTU, Modbus IP, SNMP and BACnet as what CDUs speak today, with Redfish arriving through the BMC.

SystemWhat it holds that the decision needsWhat it does not hold
Server BMCThe tray's own leak and drip sensors, the event that powered the node off, and on Lenovo's platform the blocked power permissionWhether the coolant came from this tray or ran down from the one above
CDU controllerIts internal leak tape, the level sensor, make-up pump runtime, differential pressure across the pumps, and whether it is set to alarm or to shut downWhich rack on its loop is leaking, or how much load that rack carries
BMS / BASEvery leak alarm in one place, the under-floor ropes with their zone map, and the solenoid valves it can drive if any were fittedThe severity class, or whether the row can stay up without this rack
DCIMThe rack's kilowatts, which CDU feeds which rack, and how many racks share the loopThe state of the rack's manifold valves
CMMS / EAMThe FRU tiers, the OEM service contract, who is certified on this enclosure and this CDUThe live sensor state, or the time since the cords came out
Five inputs in five systems. Which layer the leak belongs to is the join, and it is made by one person, from memory, against a clock.

That join is the interval between the alarm and the dispatch, and for a cold-plate leak it has a short and checkable form. The severity class from the matrix, the isolation point from the rack's own valves, the rack's load and its loop-mates from DCIM, and the person's class and OEM tier from the CMMS.

How to write the ten-minute leak procedure for one rack this week

The procedure below is the manuals reordered as a decision, and it fits on one page per rack type as a MOP. It produces a yes or a no on three questions before anyone touches a valve: which layer, which valve, and who.

  • Map every leak sensor on the rack to a severity class before the first alarm. The tray ropes and drip trays are major, the under-floor rope zones and the CDU's external tapes are critical, and the CDU's level and pressure alarms alone are minor until a direct sensor agrees.
  • Write the BMC's behaviour down for each platform you run. Lenovo's core module powers off and blocks power permission, NVIDIA's rack family can cut the power shelf and close valves in both directions, and a platform whose BMC only raises an event needs a person or a rack manager to do the rest.
  • Find the rack's manifold isolation valves and label them. If a rack has no valves of its own, the loop has no way to dry one rack without drying all of them, and fitting them is the first job.
  • Check the CDU's leak configuration against what you actually want. Vertiv's default is alarm only with a ten-second delay, and if you choose alarm plus shutdown, confirm thermal throttling is enabled on every cold-plated chip on that loop first.
  • Put the power-cord step first in the physical sequence. Lenovo's rule is cords out of every power conversion station before any tray is removed, and out for at least two minutes.
  • Assign by class and tier, not by proximity. A skilled or instructed person for the valve and the cords, the OEM's technician for the water loop, the CDU vendor for the unit, and name the escalation for each.
  • Apply the redundant-sensor rule to anything automatic. One sensor alarms a person, two sensors in agreement may close a valve, and neither should ever stop the CDU without a safe shutdown of the servers first.
  • Decide the load question in advance from DCIM. If this rack's neighbours can carry its work, the rack goes dark and the loop stays up, and if they cannot, that is the decision the on-call engineer needs a named person to approve.

What closes the incident is physical, not administrative. The replaced loop holds pressure, the enclosure powers on without shutting itself down, the CDU's level has stopped falling, and the four-wire rope reads dry along its whole length, and until all four are true the rack is still a dark rack on a loop that may still be losing water.

Decisionslast 6 h
Dispatch CW-2184 → M. OkonkwoWithin approved scope
Escalate CW-2168 to hall leadApproved by J. Reyes · 02:41
Close CW-2147Approved by J. Reyes · 01:12
Raise CW-2179 as WatchWithin approved scope
Reopen CW-2139 — recurrenceWithin approved scope
Extend window on CW-2151Approved by A. Mensah · 00:36
Hold dispatch CW-2133Declined by J. Reyes · 23:58
ConnectionsRead-only
BMS · Niagara N44,812 pointsRead-only
DCIM · asset + capacity1,140 assetsRead-only
Controls · CDU gateway306 pointsRead-only
CMMS · work historyWork ordersRead · write work
Control authority
Setpoint changesNot requested · not available
Start / stop equipmentNot requested · not available
Valve and damper commandsNot requested · not available
Alarm suppressionNot requested · not available
The boundary, drawn: the layer reads the sensors and the manuals, and a named person closes the valve

What this cannot do

A sequence is not a time figure for your loop. The 23 seconds is one experiment on medium-power processors with a 45 °C set point, the enclosure behaviours are Lenovo's and NVIDIA's for their own platforms, and the alarm defaults are one Vertiv CDU's, so your first ten minutes depend on which BMC, which valves and which CDU you actually run.

Nothing here takes control authority over any of it. A decision layer that reads the sensors and the manuals can say which class of leak this is and which valve and which person it belongs to, and it does not close a valve, stop a pump, drop a power shelf or trip a breaker. A named person isolates the rack, a named person pulls the cords, a named person approves any CDU shutdown, and the layer records what they decided and whether the loop's level and the rope agreed afterwards.

Answered

What happens first when a direct-to-chip cooling leak is detected?

The server acts before anyone does. On Lenovo's Neptune platforms the tray's leak sensor makes the BMC power the node off and block power permission until the leak is resolved, and NVIDIA's rack family can power off the leak tray over Redfish. The OCP paper's definition is that node-level actions belong to the BMC, rack-level actions to the rack manager or rack CDU, and row-level actions to the row CDU and the BMS.

Should you shut down the CDU when a coolant leak is detected?

Not first, and not without a safe shutdown of the servers. The OCP paper calls stopping flow at the CDU the extreme response, because every liquid-cooled server on that loop must then be shut down, and a row CDU often feeds ten or more racks. A fully loaded processor began throttling 23 seconds after loss of flow in the experiment ASHRAE cites, so the CDU is the last layer, not the first.

Where do you isolate a leaking liquid-cooled rack?

At the rack's own manifold supply and return valves, which dry one rack and leave the rest of the loop flowing. A manual ball valve isolates only one side of the network and the OCP paper warns that opening an un-terminated valve is a catastrophic leak, so the valves need labelling and a named person. Vertiv's CDU manual recommends a monitored detection system wired to close exactly these field-installed valves automatically.

Do you have to power down the rack before opening a leaking tray?

Yes, on Lenovo's enclosures. The N1380 guide requires the power cords disconnected from every power conversion station before any component is removed to inspect a leak, and left out for at least two minutes so the stations do not latch. Its tray procedure then removes the reported tray and the trays on either side of it, and a leaking manifold or quick connect is discarded and replaced rather than resealed.

How much coolant does a quick disconnect lose when you break it?

Very little, by specification. The OCP paper states that a UQD2 dry-break coupling is specified to drip under 0.063 cc when disconnected at 200 psi, and one vendor's sheet for OCP-pattern couplings gives maximum fluid loss from 0.007 cc on the smallest size to 0.03 cc on the half-inch size. The exception the paper names is a valve held open by contaminant, which is why flushing before hardware goes in matters.

Who is qualified to respond to a direct-to-chip cooling leak?

Trained personnel as defined by IEC 62368-1, in Lenovo's words, in a restricted access location the site controls. That means a skilled or instructed person for the valve and the power cords, not an ordinary user. The OEM adds its own tier: the water loop, manifolds and leakage sensor are field replaceable units that only trained service technicians install, and the CDU vendor restricts work inside its unit to trained and qualified personnel.

Can a leak detection rope miss a coolant leak?

Yes, in two documented ways. The OCP paper reports a national laboratory whose failing O-rings dripped for a long time without ever triggering the facility leak system, and it notes that slow glycol leaks evaporate and leave residue a rope will not catch. Lenovo's guide says the same for its enclosures, that a small leak may not reach either sensor and visual confirmation may be required, which is why indirect detection at the CDU matters.

Qualified responseNearest is not the same as allowed.