MTBF on cooling equipment: what it predicts, what it does not, and why the failure count decides the number
Every page on mean time between failures gives the formula and a round-number example. None says what goes in the denominator, and on a cooling plant that is decided at incident closure, by a person, and it can double the number.
By Jeel Patel, Founder at HVAC Software
The chiller's asset record says 68,000 hours between failures. It tripped on high condenser pressure in February, was repaired and returned to duty, and tripped on the same fault two days later, so the CMMS holds two work orders. In November the reliability report will count that as one failure or two, and nobody has yet decided which.
MTBF, mean time between failures, is the operating time of a repairable item divided by the number of failures in that time. On cooling equipment it predicts how many failures a fleet of chillers, CRAH units or pumps will produce in a year, not when any one unit will fail and not how long it will last. Only about 37 percent of units survive to their own MTBF, and the number moves with the rule that decides what counts as a failure. The same chiller log gives 1,000 or 2,000 hours depending on whether a fault that returns two days after the repair is counted twice or once.
- MTBF is a fleet statistic. A 68,776-hour MTBF on twelve water-cooled chillers means about one and a half failures a year across the twelve, and it says nothing about which chiller or which month.
- MTBF is not a lifetime. ASHRAE's service-life database records a median of 18 years in service for a centrifugal chiller, while the one published field failure rate for a water-cooled chiller implies an MTBF under eight years, because MTBF counts repairs and service life counts replacement.
- The denominator is a decision. Schneider's white paper on comparing MTBF figures lists twelve questions a vendor answers before counting a failure, and one of them is whether a recurring failure counts once or every time.
- The counting rule moves MTBF and MTTR in opposite directions and leaves availability alone, which is why availability is the honest number and MTBF is the one that needs its rule written beside it.
What is MTBF, and what does it measure on cooling equipment?
Mean time between failures is the predicted elapsed time between inherent failures of a mechanical or electronic system during normal system operation, and the term belongs to repairable items. A chiller, a CRAH unit or a pump is repaired and returned to service, so it has an MTBF. A fan bearing or a compressor is replaced, so it has a mean time to failure, MTTF, and the IEC's vocabulary keeps the two apart at entry 192-05-13, which defines mean operating time between failures as the expectation of the operating time between failures and reserves it for repairable items.
The maintenance profession's own definition is the same fraction. SMRP's metric 3.5.1, harmonised with the European federation's T17, is operating time divided by the number of failures, and the word that matters is operating. A standby chiller that ran 2,000 of the year's 8,760 hours has 2,000 hours in the numerator, not 8,760, and a plant that uses calendar hours for a lead-lag pair overstates the MTBF of the lag unit by whatever share of the year it sat idle.
- MTBF
- Operating time divided by failures, for a repairable item. Hours. The reciprocal of the failure rate when that rate is constant.
- MTTF
- The same fraction for an item that is replaced rather than repaired: a bearing, a compressor, a control valve. Cheung and Wang's field table uses MTTF at component level and builds the chiller's rate from it.
- MTTR
- Total repair time divided by repair events, SMRP metric 3.5.2. Moves availability and does not move reliability.
- Availability
- MTBF divided by MTBF plus MTTR. The share of time the unit is in its normal state, and the quantity the tier and redundancy models are built on.
- Annual failure rate
- Failures per unit per year. For a continuously running unit, one minus e to the power of minus 8,766 over MTBF. The number a fleet owner can plan spares against.
- L10 life
- The hours until 10 percent of a population of bearings has failed from fatigue, per ISO 281. A wear-out figure, and the one fan manufacturers publish beside an MTBF that runs to millions of hours.
- Service life
- The median years an equipment class stays in service before it is replaced, as recorded in ASHRAE's database. A replacement figure, not a failure figure.
Two things are not failures, and the definition says so. Failures that can be left unrepaired without taking the system out of service are not counted, and scheduled maintenance is excluded from the operating time, so a CRAH filter change on the preventive maintenance calendar is neither a failure nor operating time. An alarm is not a failure either, and a plant that counts every rationalised alarm as one has computed something, but not MTBF.
Why does a 69,000-hour MTBF not mean the chiller lasts eight years?
Because MTBF is the mean of an exponential distribution, and the exponential is front-loaded. Under a constant failure rate the reliability at time t is e to the power of minus t over MTBF, so at t equal to the MTBF itself the survival probability is e to the minus one, or 36.8 percent. Nearly two thirds of a population fail before reaching their own MTBF, and the distribution is memoryless, so a unit that has run 60,000 hours has exactly the same chance of failing next month as a new one.
The population arithmetic is the part that gets lost. AutomationDirect's engineering note puts it plainly: an MTBF of 40,000 hours for one device becomes 40,000 divided by two for two devices, and 40,000 divided by four for four. Twelve chillers at 68,776 hours each is a fleet MTBF of 5,731 hours, which is a failure somewhere in the plant roughly every eight months.
An MTBF of 68,776 hours is a statement about the twelve chillers together. It is not a statement about the one you are standing next to.
Schneider's White Paper 78 on MTBF makes the same point with people rather than pumps. Take 500,000 twenty-five-year-olds, watch them for a year, and record 625 deaths: the failure rate is 0.125 percent a year and the MTBF is 800 years, while the life expectancy of the same population is 75 to 80. Both numbers are true of the same people, and the MTBF is the one that says nothing about how long any of them will live.
The reason is the bathtub curve. MTBF describes the flat middle of it, the useful-life region where the failure rate is roughly constant, and AutomationDirect's note states that the rates from the electronics handbooks apply to that region and to that region only, adding that a product with an MTBF of ten years can still wear out in two. Wear-out is a rising failure rate, which in Weibull terms is a shape parameter above one, and ReliaSoft's reference says populations with β>1 have a failure rate that increases with time, which no constant-rate MTBF can see coming.
Fan manufacturers publish the wear-out number separately for exactly this reason. ebm-papst runs fans in cabinets at 40 °C and 70 °C until they fail, deems a fan failed when it deviates from its defined air flow and speed values, or when the operating noise becomes noticeable, fits a Weibull distribution and states an L10 at 40 °C and at maximum temperature, with catalogue rows such as 85,000 and 42,500 hours. Its own engineers write that the service life of the grease is very often the limiting factor for the bearing and so for the fan, and Sanyo Denki's training module gives an MTBF exceeding three million hours for many of its fans while telling the reader not to use it to judge fan life.
Where do MTBF figures for chillers, towers, pumps and CRAH units come from?
There are three origins, and a figure is only as good as the one it came from. Prediction standards such as MIL-HDBK-217F, last revised in 1995, and Telcordia SR-332 estimate a failure rate from a parts list, and they cover electronics, which is the controller in a CRAH, not the fan or the coil. Vendor field data is the second origin, and Schneider's White Paper 112 says the field data measurement method is a more accurate measure of failure rate than simulations and should always be used where a product has a field population. The third origin is your own CMMS, and it is the only one that measures your plant.
For cooling equipment as a class there is one published field set in the reliability literature. Cheung and Wang's 2019 study of water-cooled data-center cooling systems tabulates component mean time to failure and repair from on-site observations, citing Hale and Arno's 2000 survey of reliability information for HVAC components in commercial, industrial and utility installations for the chiller, tower, valve and control-panel rates and a 2012 study of pumps in high-rise residential buildings for the pump rates. The paper then builds each piece of equipment as its components in series, which gives the assembled rates in the table.
| Equipment | Failure rate, per hour (Cheung and Wang, Table 7) | MTBF implied, 1 ÷ rate | Annual failure rate implied | Reliability after one year (the paper) | Median years in service (ASHRAE database, records) |
|---|---|---|---|---|---|
| Water-cooled chiller | 1.454 × 10⁻⁵ | 68,776 h, 7.8 years | 12.0 percent | 88.04 percent | 18.0, centrifugal (239) |
| Air-cooled chiller | 1.340 × 10⁻⁵ | 74,627 h, 8.5 years | 11.1 percent | 88.92 percent | 28.5, air-cooled reciprocating (16) |
| Cooling tower | 1.483 × 10⁻⁵ | 67,431 h, 7.7 years | 12.2 percent | 87.82 percent | 17.5, metal (188) |
| CRAH | 9.609 × 10⁻⁶ | 104,069 h, 11.9 years | 8.1 percent | 91.93 percent | 20.5, computer room AC, chilled water (268) |
| Variable-speed chilled-water pump | 9.667 × 10⁻⁶ | 103,445 h, 11.8 years | 8.1 percent | 91.88 percent | 21.0, split-case single stage (133) |
| Single-speed cooling-water pump | 7.148 × 10⁻⁶ | 139,899 h, 16.0 years | 6.1 percent | 93.93 percent | 20.0, close-coupled end-suction (89) |
Read the chiller row across. A failure rate of 1.454 × 10⁻⁵ per hour is an MTBF under eight years, and the same equipment class sits in ASHRAE's database with a median of 18 years in service across 239 centrifugal chillers. The two are not in conflict, because MTBF counts every repair over the unit's life and service life counts the one replacement, and a chiller that is repaired twice in its second decade has a lower MTBF and the same service life.
Two cautions travel with the table. The rates were surveyed in commercial and industrial installations and, for the pumps, in residential towers, not measured in data halls, and the series model assumes every component failure takes the unit down. IEEE 493, the Gold Book, exists because the lack of credible data concerning equipment reliability had hindered exactly this kind of study, and the honest position is that a data-center operator's own count, kept under one rule, is worth more than any published figure.
How does the failure count change the MTBF for the same chiller?
Schneider's White Paper 112 is the one document on the open web that takes the denominator seriously, and it was written because vendors were quoting incomparable numbers. It lists the questions a vendor answers before counting: whether customer misapplication counts, whether a technician-caused load drop counts, whether a consumable that wore out early counts, whether shipping damage counts, whether a cascading failure counts once or per system, and whether recurring failures are counted multiple times or only once. Its own worked example takes one population of 2,000 units through two definitions and gets 898,462 hours with nine failures and 2,021,538 hours with four, a difference it puts at 125 percent, and it warns that comparisons made without the definitions should expect variations of 500 percent or more.
A cooling plant does the same thing to itself without noticing. The log below is a composite, one water-cooled chiller over twelve months, on duty for 6,000 of the year's hours by its run-hour counter.
| Date | Event | Loss of cooling function | Downtime |
|---|---|---|---|
| 3 February | High condenser pressure trip. Tubes cleaned, work order closed. | Yes | 4 h |
| 5 February | Same trip, 48 hours later. Fouling found on the second pass. | Yes | 6 h |
| 14 May | Oil pressure alarm. Unit kept running, sensor replaced. | No | 0 h |
| 21 August | Low refrigerant pressure trip. Leak repaired at one joint, work order closed. | Yes | 5 h |
| 28 August | Trip returns. Second joint found leaking. | Yes | 9 h |
| 2 November | Variable-speed drive fault. Drive replaced. | Yes | 6 h |
Now count it three ways. Rule A counts every work order, which is what a CMMS report does by default. Rule B counts only events that took the cooling away, which is Schneider's Type I definition. Rule C is Rule B with a recurrence rule, in which a fault that returns inside a seven-day observation window reopens the failure it belongs to rather than starting a new one.
| Counting rule | Failures | MTBF | MTTR | Availability |
|---|---|---|---|---|
| A. Every work order is a failure | 6 | 6,000 ÷ 6 = 1,000 h | 30 ÷ 6 = 5.0 h | 99.50 percent |
| B. Only loss of cooling function counts, every trip counts | 5 | 6,000 ÷ 5 = 1,200 h | 30 ÷ 5 = 6.0 h | 99.50 percent |
| C. Loss of function, and a return inside seven days reopens the same failure | 3 | 6,000 ÷ 3 = 2,000 h | 30 ÷ 3 = 10.0 h | 99.50 percent |
That last column is the finding. Total operating time and total downtime are physical facts recorded by the BMS, so availability is fixed by the year, while the number of failures is a decision and MTBF and MTTR are two fractions with that decision in the denominator. A site that changes its recurrence rule in March will report an MTBF improvement in December that no chiller earned.
The counting rule cannot make the plant more available. It can only make the MTBF look better and the MTTR look worse, by exactly the same factor.
The node that decides the number is the one no reliability page names. Whether the 5 February trip is the 3 February failure coming back or a new failure is a question about evidence, about whether the condition had actually cleared before the work order was closed, and that is the question the verified-recovery argument exists to answer. This page only shows what the answer does to the metric.
What does MTBF predict for a cooling plant, and what does it not?
| MTBF predicts | MTBF does not predict |
|---|---|
| How many failures a fleet will produce in a year: units × 8,766 ÷ MTBF for continuously running equipment | When any one unit will fail. Under the exponential model a unit that has run for years is no more likely to fail next month than a new one |
| The spares and call-out budget for the year, once the count is under one rule | When wear-out begins. That is an L10 or a Weibull shape parameter above one, and the flat-region MTBF cannot see it |
| The failure rate a redundancy model needs: Cheung and Wang's N+1 and 2N results are built on the rates in the table above | How long a failure leaves the hall. That is thermal ride-through, and it depends on what failed and what is stored, which is ranked in its own guide |
| A comparison between two years, or two sites, that count under the same written rule | A comparison between two vendors, or two sites, that do not. White Paper 112's 500 percent applies |
| With MTTR, the availability the tier and redundancy targets are stated in | Anything about the failure that just happened. It is a mean over a population and a period, not a diagnosis |
The left column is where the number earns its keep. Cheung and Wang feed their component rates into a plant model and find that one redundant chiller lifts a water-cooled plant's availability from 99.679 to 99.883 percent, that N+2 and 2(N+1) add only 0.0001 and 0.0003 percent over N+1 and 2N, and that a second chiller plant on a different cooling source can double the plant's one-year reliability. None of those answers exists without a failure rate, and every one of them is wrong by the same factor as the count that produced it.
The right column is where cooling plants get hurt. Uptime Institute's outage analysis for 2023 reports that on-site power problems remain the biggest cause of significant site outages and that cooling failures, software errors and network issues stand out as particularly troubling behind them, and its 2025 report finds that the failure of staff to follow procedures has become an even greater cause of outages than the year before. A failure caused by a skipped procedure is a failure of the plant whether or not the MTBF rule counts it, and White Paper 112 notes that vendors routinely filter human causes out.
How to calculate MTBF for your cooling fleet this week
The formula takes a minute. The rule takes a meeting, and the meeting is the work. The order below produces a number that can be compared with next year's, which is the only comparison MTBF is good for.
- Define the population before the period. All CRAH units in hall B, or all water-cooled chillers on the plant, of similar design and duty. A chiller and a CRAH in one MTBF is a number about nothing.
- Take operating time from the BMS run-hour counters, not the calendar. A lead-lag pair has two different numerators, and a standby unit's idle hours are not operating time.
- Write the failure definition down, in Schneider's terms. Type I, the unit stopped doing its job, is the one that matters for cooling. Decide whether human-caused failures count, and say so.
- Set the recurrence rule and the observation window per equipment type. A return inside the window reopens the same failure. This is the rule that moved the composite chiller from 1,000 to 2,000 hours, and it belongs on the same page as the number.
- Pull the work orders for the period, remove the scheduled maintenance, and count under the rule. Keep the list of what was excluded and why, because next year's reviewer will ask.
- Compute MTBF as operating hours divided by failures, then the annual failure rate as one minus e to the power of minus 8,766 over MTBF, then the fleet expectation as units times 8,766 divided by MTBF. The last one is the number to budget from.
- Record the rule beside the number, with the population, the period and the excluded events. An MTBF without its rule is a figure without units.
The arithmetic for a hall, using the published CRAH rate above as the input: twenty CRAH units at 9.609 × 10⁻⁶ per hour, running continuously, is 20 × 8,766 × 9.609 × 10⁻⁶, or about 1.7 failures a year across the twenty. Twelve water-cooled chillers at 1.454 × 10⁻⁵ is about 1.5 a year. Those are budgets for spares and call-outs, and they are also the expected failure counts the next year's own log will be judged against, under the same rule.
What do the BMS, DCIM and CMMS already hold?
Each system holds one term of the fraction and none holds the rule. The gap is not that the systems fail to talk. It is that the numerator lives in one, the denominator in another, and the recurrence decision in neither.
| System | What it holds about the count | What it does not hold |
|---|---|---|
| BMS or BAS | The run-hour counters, the trips and the return-to-normal timestamps, which are the operating time and the downtime | Which trip belongs to which failure, or whether the closed work order's condition had actually cleared |
| DCIM | Which units are duty, which are standby, and what the redundancy state was when the unit tripped | Any failure count. It is the model the count is fed into |
| CMMS or EAM | The work orders, their closure times and the MTBF report built on them, one row per work order by default | The recurrence rule. A returning fault is a new work order unless somebody joins it to the last one by hand |
| The vendor's MTBF | A field-data or predicted figure for the product population, under the vendor's own definition of failure | Your plant, your duty cycle or your rule. White Paper 112's 500 percent lives here |
| The commissioning record | The first run hours and the first measured restart for every unit, taken under load with calibrated instruments | Anything after handover, unless the readings were handed over as values rather than a report |
The join is possible today, in pieces. Maintenance Connection's integration with Accruent Observe lets telemetry data verify whether the issue has been fixed before allowing the status change, which is one CMMS gating one closure on one point. The starting record every unit already has is the commissioning baseline, the first run hours and the first measured restart, and a count that begins there and runs under one rule is the only MTBF a plant can trust against itself.
Who decides what counts as a failure?
A named person does, before the work starts, and the decision is the recurrence rule and the observation window per equipment type. The technician closes the work order and owns the finding at the unit. The reliability lead owns the rule and the count, and the rule is written down where the number is reported, so that a change to it is visible as a change.
What the layer would do is read. It would hold the operating hours from the BMS and the closures from the CMMS against the rule, flag the 5 February trip as inside the 3 February window, and report both the count and the rule that produced it. It would not close the work order, reopen it, or change its status, and it would not start, stop, stage or reset the chiller.
What this cannot do
A page assembled from a reliability paper, two survey citations and ASHRAE's database is not your plant's failure rate. The rates in the table were surveyed in commercial and industrial buildings and residential pump rooms and assembled in series by one study, the service-life medians are what 345 buildings reported to a database, and the composite chiller is a composite. Nothing here is a benchmark to be measured against, and the only benchmark that exists is last year's count under this year's rule.
Nothing here predicts a failure or diagnoses one. MTBF is a mean over a population and a period, and a decision layer that holds the count can say that this hall produced four Type I failures against an expectation of 1.7 and that two of them were returns inside the window. It cannot say which CRAH fails next, and it does not start, stop, stage or reset anything. A named person sets the rule, closes the work, and owns the number.
Answered
What is a good MTBF for a chiller?
There is no published benchmark, and the one field failure rate in the reliability literature implies an MTBF of about 68,776 hours for a water-cooled chiller, surveyed outside data halls and assembled in series by one study. A good MTBF is one computed under a written failure definition and recurrence rule, compared only with the same plant's earlier years under the same rule. A vendor's figure is not comparable without its definition.
Is MTBF the same as lifespan?
No. MTBF is the operating time between repairs across a population in its constant-failure-rate years, and lifespan or service life is when the unit is replaced. ASHRAE's database records a median of 18 years in service for a centrifugal chiller while the published field rate implies an MTBF under eight years, because the chiller is repaired several times before retirement. Fans show the gap more sharply, with L10 lives in tens of thousands of hours beside MTBFs in millions.
What is the difference between MTBF and MTTF?
MTBF applies to repairable items and MTTF to items that are replaced. A chiller, a CRAH or a pump has an MTBF because it is repaired and returned to duty. A fan bearing, a compressor or a control valve has an MTTF, and Cheung and Wang's model gives the chiller its failure rate by adding the failure rates of its components. The IEC's vocabulary keeps the two terms apart for that reason.
Which is better, MTBF or MTTR?
Neither on its own, and availability needs both. MTBF divided by MTBF plus MTTR is the availability the tier and redundancy targets are stated in. On a cooling plant the counting rule moves MTBF up and MTTR up together, by the same factor, and leaves availability alone, so availability is the number the rule cannot flatter. MTTR also hides the alarm-to-dispatch interval, which has its own page.
How do I convert MTBF to reliability?
Under a constant failure rate, reliability over a period t is e to the power of minus t over MTBF. At t equal to the MTBF that is 36.8 percent, at half the MTBF it is 60.7 percent, and over one year of 8,766 hours a chiller with an MTBF of 68,776 hours has a reliability of about 88 percent, which is the figure Cheung and Wang report for one year of operation. The conversion is only valid before wear-out.
Does N+1 change the MTBF of the cooling plant?
It changes the plant's availability, not each unit's MTBF. Cheung and Wang's model takes the same unit failure rates and finds that one redundant chiller raises a water-cooled plant from 99.679 to 99.883 percent availability, while N+2 and 2(N+1) add only 0.0001 and 0.0003 percent over N+1 and 2N. The units fail at the same rate in every configuration. The redundancy decides how often the hall notices.
Should a failure that comes back after the repair count twice?
It should count under whatever rule the site wrote down before the work started, and that rule should be the recurrence rule and observation window used to close incidents. Schneider's white paper lists it as one of the questions vendors answer differently, and the composite chiller on this page moves from 1,000 to 2,000 hours on it alone. A count that changes its rule mid-year is not a trend.