CRAC unit not cooling: the diagnostic sequence, keyed to the alarm the unit has already raised
Every troubleshooting page lists causes. None of them starts where the unit does: with the alarm its controller raised, the protective action it has already taken, and the readings the building management system was showing before anyone picked up the phone.
By Jeel Patel, Founder at HVAC Software
CRAC-04 raised a high temperature warning at 02:10. Return air reads 84 °F against a 75 °F setpoint, the fan is running, and the display shows High Head Pressure C1 acknowledged twice since midnight. The on-call engineer is forty minutes out and the question on the phone is whether anyone should press reset.
A CRAC unit that is not cooling is in one of three states, and the alarm on its controller says which. Either the unit has stopped on loss of power, main fan overload or loss of airflow, or the fan is running while the controller has locked the compressor off on high head pressure, low suction pressure or short cycling, or everything is running and the supply air is still warm because of a setpoint, a load beyond capacity, dirty filters, low charge or a valve. The sequence starts from that alarm, reads what the building management system already shows, and only then sends a person to the unit with the manufacturer's cause list in the manufacturer's order, inside ASHRAE's fifteen-minute window.
- The controller has already acted before anyone reads the alarm. On Liebert iCOM as documented, a high head pressure alarm has removed the compressor, a loss of airflow alarm has stopped a DX unit, and a high temperature alarm has done nothing at all, because it is a warning about a reading rather than a fault.
- The BMS can put the unit in one of the three states without a site visit. Fan status, compressor status per circuit, return and supply air temperature and the filter differential are enough to know whether the walk is to a stopped unit, a locked-out compressor or a unit that is simply behind its load.
- The check order is the manual's, not a generic pressure guide's. Each branch below follows the Liebert DS and Challenger troubleshooting tables, and the line between what an operator may check and what needs a certified technician is drawn by 40 CFR 82.161.
- Three clocks run at once and they are not the same clock. ASHRAE's 5 °C in any fifteen minutes is the hall's, the compressor's own delays and lockouts are the unit's, and the alarm counter that holds a compressor off until a person resets it is the one that decides whether a unit is actually back.
What is a CRAC unit, and what does "not cooling" mean on one?
A computer room air conditioner (CRAC) is a direct-expansion cooling unit with its own compressor and refrigerant coil, rejecting heat through an air-cooled condenser, a drycooler or a water loop. Vertiv's Liebert DS guide covers that family from 35 to 105 kW in air-cooled, water and glycol-cooled, Dual-Cool and GLYCOOL forms, and the unit controls to return air temperature by default, with a supply air sensor as an option.
"Not cooling" is three physically different conditions that produce the same phone call. The unit can have stopped, the fan can be moving air past a coil with no refrigerant flowing through it, or the whole machine can be running and losing to the load in the room.
- Return-air sensor and setpoint
- The reading the unit controls to. A high temperature alarm is this reading crossing its alarm setpoint, and nothing more.
- Supply-air sensor
- Optional on most units. Without it the BMS cannot see the cooling-unit delta-T, which is the one reading that separates a unit behind its load from a unit that has stopped cooling.
- Compressor, high-pressure switch, low-pressure switch
- The refrigeration circuit and its two guards. The switches open the contactor, and the controller decides whether that trip becomes a lockout.
- Condenser or drycooler, head-pressure control
- Where the heat goes. Fan-speed control or a flooded receiver holds head pressure in cold weather, and a fault here shows up as a high or a low head pressure trip.
- Liquid line solenoid valve and expansion valve
- The solenoid closes on a stop so the compressor can pump down. The thermostatic expansion valve meters refrigerant into the coil, and its setting is read as superheat.
- Fan and the differential pressure switch
- The switch proves the fan is moving air. When it does not, the unit raises loss of airflow and, as configured on a DX unit, shuts down.
- Filter clog switch
- A second differential switch across the filters, adjustable at the switch. It raises a warning, and a warning does not stop anything.
- Unit controller
- Holds the alarm history, the high-pressure alarm counters and the safety logic. Its actions run whether or not the alarm was classified, enabled or acknowledged.
Why does the unit's own alarm decide where the sequence starts?
The controller acts before anyone reads the display. On Liebert iCOM as documented, an event with a safety function such as high pressure, low pressure or main fan overload executes that function in any case, independent of whether the event was set to alarm, warning or message, and independent of whether it was enabled at all.
That is why the alarm name carries the state. A high temperature event on that controller is fixed to a warning, ignored for the first minute after the fan starts and delayed thirty seconds by default, and it changes nothing about the machine. A high head pressure event has already removed a compressor by the time it is annunciated, and the compressor stays off until a person resets the counter.
| Alarm | What triggered it | What the controller has already done | What that already tells you | Where the sequence starts |
|---|---|---|---|---|
| High Temperature | Return air reached the high temperature alarm setpoint | Nothing. It is a warning, held back for one minute after fan start | The unit is running. Either its capacity is short of the load or a cooling component has failed without its own alarm | Branch C, after checking whether a second alarm is present |
| High Head Pressure | The head pressure switch opened, above 360 psig on the Challenger | Opened the compressor contactor. Silent retries in the first ten minutes of a run, immediate lockout after that, and a lockout after three trips in a rolling twelve hours. Stays off until the HP alarm counter is reset to zero, even if pressure has fallen | Heat is not leaving the condenser side. The fan may still be running, so the room is being stirred, not cooled | Branch B, on the condenser or the fluid loop |
| Low Suction Pressure | The suction switch opened below its factory preset. The input is ignored for the low pressure delay, usually three minutes on air-cooled units, then must stay open five minutes | Stopped the compressor on the switch. The alarm stays active ninety minutes after acknowledgement | Refrigerant is short at the coil: a leak, a closed valve, a stuck solenoid, a starved expansion valve, or filters so dirty there is no load on the coil | Branch B, on the refrigerant side |
| Short Cycle | More than ten cooling starts in one hour, or five low-pressure cycles in ten minutes | Logged it. The three-minute short-cycle delay is already holding the compressor between starts | Either the charge is low but not low enough to trip low suction, or the room load is small against the unit's capacity | Its own question, which the short-cycling guide will own |
| Loss of Airflow | The differential pressure switch across the fan did not prove air after a three-second default delay | Shutdown, where that option is configured, which the manual says is intended for DX models | The unit has stopped. A blocked inlet or outlet, a broken belt, a loose wheel, or a fan contactor | Branch A |
| Main Fan Overload | The motor protection or EC fan fault tripped after a five-second default delay | Shutdown as configured. While this alarm is active, loss of airflow is masked | The fan motor itself. On an EC motor the stop is electronic with no automatic restart | Branch A |
| Loss of Power | Power returned to a unit whose fan was on when it was lost | Raised on restoration only. Resets thirty minutes after acknowledgement | The unit may have restarted, or may be waiting on its autorestart delay, or may not have come back | Branch A if the fan is not running |
| High and Low Temperature together | The temperature input signal was lost | Displayed dashes and initiated 100 percent cooling | The sensor or its cable, not the refrigeration | Branch C, at the sensor |
One consequence is easy to miss and expensive. On the same controller, a configured standby unit is started by an event of type alarm only, so a high head pressure lockout starts the standby and a high temperature warning on a unit that is merely behind its load does not.
What does the BMS show before anyone walks to the unit?
The unit's common alarm relay and its points are almost always already on the BMS, and four of them place the unit in one of the three states from a desk. The iCOM alarm relay outputs include high temperature, high head pressure per compressor, loss of airflow and change filters, and the analogue points carry return air, supply air where the sensor is fitted, and fan and compressor status.
Reading them in the order below turns a call about a warm room into a call about a specific branch. The walk to the unit then starts with the right tools and, where the branch needs it, the right card.
| Reading | Unit stopped | Fan running, compressor locked off | Everything running, supply air warm |
|---|---|---|---|
| Fan status, or airflow proved | Off, or airflow not proved | On | On |
| Compressor status, per circuit | Off | Off on the locked circuit, with the high-pressure counter above zero or the low-pressure alarm active | On, or cycling |
| Return air temperature | Rising toward the room's hot-aisle temperature | Rising, more slowly, because the fan is still mixing the room | Above setpoint and steady, or climbing with the load |
| Supply air temperature and delta-T | Supply equals return within minutes: no delta-T | Supply approaches return: delta-T collapsing toward zero | Delta-T present but the supply is warm, or delta-T high and the unit flat out |
| Filter differential or clog warning | Any | Any | A clog warning here is the first cause on the low-suction list |
| Active alarm and its type | Loss of airflow, main fan overload, or loss of power | High head pressure, low suction pressure, or short cycle | High temperature only, or no alarm at all |
| The neighbouring units | One or more picking up return temperature | Same | All of them running near the same return temperature, which is a room problem rather than a unit problem |
The delta-T column needs a baseline to read against. Purkay Labs puts the ideal cooling-unit delta-T between 15 and 25 °F, Upsite puts a legacy DX CRAC near 18 °F by design, and the number that matters is the one your unit produced at commissioning at a known load.
A low delta-T with a cold supply is a different finding from a low delta-T with a warm one. The first is bypass, which Schneider's White Paper 42 describes as a return air temperature considerably below room ambient indicating a short circuit in the supply path, and Upsite states as a return more than 5 °F below the IT exhaust. The second is the unit failing to remove heat, and it belongs to branch C.
A high temperature alarm says the room is warm. A high head pressure alarm says the unit has already taken its compressor away. Those are not the same call, and the BMS knows which one it is before anyone leaves the desk.
What is the diagnostic sequence, step by step?
The sequence has two steps before any branch and one after every branch, and the branches follow the manuals' own troubleshooting tables. Each check names what it rules out, because a check that rules nothing out is a guess with a torch.
- 1. Confirm the hall before the unit. Read the rack inlet sensors downstream of the unit against the ASHRAE A1 allowable ceiling of 32 °C and the 5 °C change in any fifteen minutes, count the units still running against the load, and ask whether this failure is bounded to one unit or shared by every unit on the same condenser, drycooler or loop, which is the question the CRAC versus CRAH guide answers failure by failure.
- 2. Read the alarm, the event log and the counters before touching anything. The order of events in the log is the diagnosis half done: a clogged filter warning an hour before a low suction alarm is a different unit from a high head pressure alarm on a hot afternoon with nothing before it.
- 3. Branch A, the unit has stopped. Check main power at the disconnect, fuses and breakers to the fan, the 24 VAC control voltage, and the fan overload reset, then the inlet and outlet for blockage, belts and wheel tightness on a belt-drive fan, and the fan contactor. On an EC fan the Liebert DS guide says the motor stops electronically on overtemperature, rotor position, locked rotor or phase failure with no automatic restart, and power must be off for at least twenty seconds before it will run again.
- 4. Branch B, the fan runs and the compressor is locked off. For high head pressure on an air-cooled unit, the Challenger manual's list is power to the condenser, condenser fans, the head-pressure control valve, closed service valves, dirty condenser coils and crimped lines. On a water or glycol unit it is the regulating valve, whether the pumps are running and the service valves open, whether the tower or drycooler is running, and whether the fluid entering the condenser is at or below design temperature. For low suction it is loss of refrigerant through leaks or crimped lines, then the liquid line solenoid, the low-pressure switch, the expansion valve and the head-pressure control valve, then closed service valves at the liquid line, condenser or receiver. Reset the counter only when the cause is found, because the reset restarts the compressor.
- 5. Branch C, everything runs and the supply air is warm. Check the setpoints and where the return sensor actually sits, then the load against capacity, since one watt consumed needs one watt of cooling and a unit behind its load is not a broken unit. Then filters, then the charge by subcooling and the expansion valve by superheat with gauges on the circuit, then a plugged filter-drier, then under-floor restrictions and closed dampers near the unit, and after any service, compressor rotation.
- 6. Close on a reading, not a reset. Supply air back at setpoint, the unit's delta-T back to its commissioned value at the current load, the counters at zero with the cause written against them, and the standby returned to standby.
| Step | What you check | What it rules out | Who may do it |
|---|---|---|---|
| 1. The hall | Rack inlets against the ASHRAE window, units running against load, shared or bounded failure | Whether this is a unit problem or a plant problem, and whether there is time for a diagnosis at all | The operator, from the BMS and DCIM |
| 2. The log | Active alarm, event order, high-pressure counters, low-pressure alarm state | Two of the three states, before anyone moves | The operator, at the display or the BMS |
| 3A. Power and fan | Disconnect, fuses, control voltage, overload reset, blockage, belts, contactor, EC fan fault | Every refrigeration cause, because a stopped fan is not a refrigeration problem | The operator for the reads and the reset. A qualified technician to open the EC motor, per the DS guide |
| 3B. Condenser side | Condenser power and fans, head-pressure control, service valves, coil cleanliness, fluid flow and entering temperature | Charge and coil causes, if the heat-rejection side is proven | The operator for status and valves. A certified technician for anything on the circuit |
| 3B. Refrigerant side | Leaks, solenoid, low-pressure switch, expansion valve, receiver and liquid-line valves, the 32 psig floor the DS compressor needs to run | Everything downstream of the charge | Type II or Universal under 40 CFR 82.161 |
| 3C. Load and air | Setpoints, sensor location, load against capacity, filters, under-floor restrictions, bypass | A refrigeration fault, if the air side alone explains the supply temperature | The operator and the facilities team |
| 3C. Charge and valve | Subcooling against outdoor ambient, superheat at the expansion valve bulb, filter-drier temperature drop | Airflow causes, if the readings are off with clean filters and proven air | Type II or Universal, with gauges on the circuit |
| 4. Closure | Supply air, delta-T at baseline, counters at zero, cause recorded, standby returned | The temporary recovery that comes back the next hot afternoon | The operator, with the technician's finding attached |
The card is the line in that last column, and it is a legal line rather than a competence one. Under 40 CFR 82.161, anyone who could reasonably be expected to violate the integrity of the refrigerant circuit must be certified, Type II for high-pressure appliances, which puts gauges, a charge, a superheat adjustment and a compressor swap on one side of the line and a filter change, a valve position check and a display reset on the other. The technician qualification guide carries the full ledger.
Two figures from the manuals belong in the technician's pocket for branch C. The Challenger table says to reset the expansion valve for 10 to 15 °F of superheat and the older DS user manual gives 10 to 20 °F, while the 2026 DS guide prints the setting as a blank and says only that too little refrigerant fed to the coil gives high superheat, too much gives low, one turn at a time, fifteen minutes loaded before rechecking.
How long does the hall have while you diagnose?
The hall's clock is ASHRAE's. Table 2.1 of the 2021 thermal guidelines puts the recommended envelope at 18 to 27 °C and the A1 allowable at 15 to 32 °C, and note f limits the change to 20 °C in an hour and no more than 5 °C in any fifteen-minute period, which the standard says is a change within a window and not a rate.
How fast that window closes depends on density, and the published figure is stark. Schneider's White Paper 179 says that without any cooling the maximum rate of rise could easily be 5 °C per minute or more depending on density and room layout, and that a DX CRAC that has stopped may take several minutes to restart once power or the reset is back.
- The hall's clock
- 5 °C in any fifteen minutes and 32 °C at the A1 inlet. Runs from the moment the unit stopped removing heat, which on a locked-out compressor is earlier than the alarm.
- The unit's clocks
- Three minutes of short-cycle delay between starts, a ten-minute window in which high head pressure retries silently, five minutes for a low-suction trip to confirm, thirty minutes of digital-scroll lockout on high temperature, ninety minutes before a software alarm clears itself.
- The reset clock
- Indefinite. A compressor locked off on the high-pressure counter stays off until a person sets the counter to zero, and the manual is explicit that falling pressure does not release it.
- The standby's clock
- Zero, if a standby is configured and the event was an alarm. Never, if the event was a warning.
The three do not add up to a single number, and this page does not compute one. What each unit type buys after a power loss, and how the DX CRAC ranks against a chilled-water hall, is the failure-time ranking's question, and whether the spare can actually carry the load is the redundancy guide's.
What do the unit controller, BMS, DCIM and CMMS already hold?
All four hold something true about the unit, and none of them holds the join between the alarm, the state it produced and the branch that follows. That is different from saying the systems do not talk to each other, and it is the honest description of the situation.
| System | What it already holds | What it does not hold |
|---|---|---|
| Unit controller | The alarm history, the event order, the high-pressure and high-temperature counters, the safety logic that already acted, and on iCOM a wellness calculation that sets time to next maintenance to zero on a pressure alarm | Anything about the room beyond its own return sensor, or which unit should carry the load instead |
| Building management system (BMS) | The common alarm relay, the analogue points, the trend of return and supply air, fan and compressor status, and the alarms it classifies and routes | Which state the unit is in, because it stores the points and not the reading of them, and the commissioned delta-T to read against |
| Data center infrastructure management (DCIM) | Hall load, rack inlet temperatures, capacity against N and which racks sit downstream of this unit | The unit's refrigeration state, or that its compressor is locked out on a counter |
| Computerised maintenance management system (CMMS) | The work order, the technician, the last filter change and the closed ticket | The reading that closed it, or whether the counter was reset with the cause written or without |
| Operating layer | Reads all four and holds the unit as an object with its baseline delta-T and its alarm history (verified now). Puts the state, the branch, the window and the credential on the work order and closes it on the supply-air reading (design-partner scope) | Any control authority. It does not reset a counter, start a compressor, cycle a fan or start a standby |
Alarm classification is its own discipline, and the warning-versus-alarm distinction the controller draws is exactly the one ISA-18.2 asks a site to rationalise deliberately. The unit ships with high temperature fixed to a warning for a reason, and a site that routes warnings as if they were lockouts, or lockouts as if they were warnings, has undone that reasoning at the BMS.
How to run this sequence from the BMS this week
None of this needs new equipment. It needs the points the unit already exposes to be on the BMS with names a person at three in the morning can read, and one baseline per unit.
- Map the unit's alarm relay outputs individually rather than as one common alarm: high temperature, high head pressure per compressor, loss of airflow and change filters are four different states, and a single common contact collapses them into one.
- Trend the high-pressure alarm counter, or where it is not exposed, the compressor status per circuit against the high head pressure event. A compressor that is off with the counter above zero is locked out, and no amount of falling pressure will start it.
- Fit or wire the supply air sensor where the unit has none, and record each unit's delta-T at a known load on a day it is clean. That number is what turns every later reading into evidence.
- Write the three states into the emergency operating procedure (EOP) as the first triage question, with the BMS reads that answer it, and the branch each state opens. The preventive maintenance checklist already carries the filter interval that removes the most common branch C cause before it becomes a call.
- Name who on the on-call rota holds a Type II or Universal card, because branch B and half of branch C cannot be completed without one, and a dispatch that discovers this at the unit has wasted the window.
- Check the teamwork configuration on every unit: which events are alarms, which standby they start, and whether a high temperature warning on a unit behind its load leaves the standby idle by design.
CRAH-07 supply air above limit
- ✓Confirm redundancy state at panel
- ✓Isolate per LOTO — CHW-B branch
- Inspect valve actuator travel
- Verify supply air returns below 24 °C
- Restore N+1 and record final state
The change that pays first is the counter. It is the one point that distinguishes a unit that will recover on its own from a unit that will sit at a warm return temperature until a person arrives, and most sites do not trend it.
What this cannot do, and what we do not claim
The alarm names the state and not the cause. A high head pressure lockout on an air-cooled unit has six causes on the manufacturer's list, a dirty coil and a dead condenser fan among them, and nothing on the BMS distinguishes them without someone at the condenser.
The product boundary is the one this site states everywhere. The operating layer reads the controller, the points and the work record, and it does not reset a counter, start a compressor, cycle an EC fan, change a setpoint or start a standby, because a person with the authorisation does that and a person signs the return to service.
Two further limits belong in the open. Every threshold here is Liebert's as documented for the Challenger, DS and iCOM families, so the manual for your unit wins where it differs, and the delta-T ranges are two named authors' figures rather than a measurement of your hall.
Which guides sit beside this one?
The failure-by-failure spread of a CRAC or CRAH, the maintenance list that prevents the most common branch, the credential ledger and the redundancy arithmetic each have their own page. This one stops at the ordered sequence for one symptom on one unit, and the alarm-to-dispatch interval is where the minutes between that unit's alarm and the right person being sent are counted.
Answered
Why is my CRAC unit running but not cooling?
Because the fan and the compressor are separate machines with separate protection. On Liebert iCOM as documented, a high head pressure trip opens the compressor contactor and locks it off after three trips in a rolling twelve hours, while the fan keeps moving air past a coil with no refrigerant flowing through it. Read compressor status per circuit and the high-pressure counter on the BMS before anyone walks over.
What does a high head pressure alarm on a CRAC unit mean?
It means the discharge pressure crossed the switch setting, above 360 psig on the Liebert Challenger, and the controller removed the compressor. On an air-cooled unit the manual's checks are power to the condenser, condenser fans, the head-pressure control valve, closed service valves, dirty condenser coils and crimped lines. On a water or glycol unit they are the regulating valve, pump and valve state, the tower or drycooler, and the entering fluid temperature.
What causes low suction pressure on a CRAC unit?
Refrigerant being short at the coil. The Challenger manual's list runs from loss of refrigerant through leaks or crimped lines to the liquid line solenoid, the low-pressure switch, the expansion valve, the head-pressure control valve and closed service valves, and its troubleshooting table adds dirty air filters, a plugged filter-drier, wrong superheat and poor air distribution. The DS guide sets 32 psig as the floor the compressor needs to run.
Does a high temperature alarm mean the CRAC unit has failed?
No. On Liebert iCOM the high temperature event is a warning fixed to that type, raised when return air reaches the alarm setpoint and held back for one minute after the fan starts. The manual's own checks are the setpoints, whether the room load exceeds the unit's capacity, and then whether the compressor and valves are operating. A unit behind its load is not a broken unit, and a warning does not start a configured standby.
How long can a data hall run without a CRAC unit?
Long enough to diagnose only if the remaining units carry the load. ASHRAE's 2021 thermal guidelines allow a Class A1 inlet up to 32 °C and no more than 5 °C of change in any fifteen minutes, and Schneider's White Paper 179 says the rate of rise without cooling could easily be 5 °C per minute or more depending on density and layout. The answer is the spare unit's, not the failed unit's.
What should the delta-T across a CRAC unit be?
Its own commissioned value at a known load. Purkay Labs gives 15 to 25 °F as the ideal cooling-unit delta-T and Upsite puts a legacy DX CRAC near 18 °F by design, so a unit reading well under its baseline with a cold supply is seeing bypass air, and one reading under its baseline with a warm supply is not removing heat. Without a supply air sensor the BMS cannot see this reading at all.
Can I reset a locked-out CRAC compressor myself?
The reset is a display action, and the manual is explicit that setting the high-pressure alarm counter to zero restarts the compressor without further confirmation. Whether you should is the question, because the cause that tripped it is still there and the next trip counts toward the lockout. Anything beyond the display that opens the refrigerant circuit needs a Type II or Universal card under 40 CFR 82.161.