Mean time between failures MTBF
MTBF is the average operating time between failures of a repairable item, computed across a population as accumulated operating hours divided by the number of failures. It is a rate turned upside down, and it is the most systematically misread number in equipment specification. A 500,000-hour MTBF is not a 57-year life, even though 500,000 hours divided by 8,760 is 57.
It is a statement that a fleet accumulating 500,000 operating hours should expect about one failure — so 100 units running for a year, which is 876,000 device-hours, would expect between one and two failures among them. Under the constant-hazard assumption that makes λ = 1/MTBF valid, the probability that any individual unit survives to t = MTBF is exp(-1), about 37 percent.
Reviewed August 2026 by Sergey Syrvachev
New to BESS? Start free with the 7-email fundamentals course — no cost, no account.
What the number counts, and what it needs to mean anything
The computation is unglamorous: total accumulated operating hours across the population, divided by the number of failures observed in those hours. Three inputs are hiding in that sentence. The item boundary — is the failing thing a module, a rack, a converter, a container?
The failure definition — does a latched trip cleared by a remote reset count, does a BMS-isolated module count, does a derate count? And the population and duration, because a figure derived from twelve failures over three million device-hours is an estimate with a confidence interval around it, while a figure derived from a calculation has no interval at all and no field behind it.
MTBF applies to repairable items, where the clock restarts after each restoration and the same serial number can contribute several failures. That is the right frame for a PCS, an HVAC skid, a container, a plant. It is the wrong frame for items you replace rather than repair — a cell, a fuse, a semiconductor die — where the corresponding metric is MTTF, mean time to failure, and each item contributes exactly one observation. Vendors mix the two words freely. The distinction decides whether the statistic describes a repair cycle or a one-way life.
The two conventions, and why you must say which one you are using
There are two definitions of the interval in circulation and they differ by the repair time. In the first, MTBF counts operating time only — the convention IEC 60050-192, the international dependability vocabulary, reflects when it names the quantity mean operating time between failures. Under that convention the steady-state identity is A = MTBF / (MTBF + MTTR), which is the form the availability entry on this site states and the form used throughout these entries.
In the second, MTBF is measured failure to failure and therefore contains the repair: MTBF = MTTF + MTTR. Under that convention the identity becomes A = MTTF / MTBF, and writing A = MTBF / (MTBF + MTTR) double-counts the downtime. Neither convention is wrong; using one and quoting the other is. When you receive an MTBF in a technical proposal, ask in writing whether restoration time is inside it — the answer changes the availability the number implies, and on equipment with long repair cycles it changes it materially.
MTBF is for repairable items and MTTF for replaced ones. Two conventions circulate — operating time only, or failure to failure — so state which. The provenance is usually a parts-count prediction rather than field data, and service life is set by wear-out, which the flat-hazard model explicitly does not describe.
- What it is
- Accumulated operating hours ÷ number of failures, across a population of repairable items
- Not a life
- 500,000 h MTBF ≠ 57-year life; under constant hazard, survival to t = MTBF is exp(-1) ≈ 37 percent
- Fleet reading
- 100 units × 1 year = 876,000 device-hours; at 500,000 h MTBF that is between one and two expected failures
- MTBF vs MTTF
- MTBF for repairable items (PCS, HVAC skid, container); MTTF for items you replace (cell, fuse, die)
- Convention A
- Operating time only — IEC 60050-192's mean operating time between failures; gives A = MTBF / (MTBF + MTTR)
- Convention B
- MTBF = MTTF + MTTR (failure to failure); gives A = MTTF / MTBF — state which you mean
- λ = 1/MTBF
- Valid only under a constant hazard rate, i.e. the flat middle of the bathtub curve
- Usual provenance
- A parts-count prediction (MIL-HDBK-217F, Telcordia SR-332 style), not field data — ask which
- Series arithmetic
- System MTBF is lower than every component MTBF in the chain
- Contract status
- LTSAs guarantee availability (commonly 95-98 percent annually), not MTBF
Vendor MTBF is usually a prediction, not a measurement
Most MTBF figures on a datasheet are calculated, not observed. The classic method is a parts-count or parts-stress prediction from a reliability handbook — MIL-HDBK-217F, whose last notice dates from the 1990s, or Telcordia SR-332 in telecom-derived practice — where each component's tabulated base rate is scaled for temperature, electrical stress and environment, and the results are summed as a series chain.
The output is arithmetic on a bill of materials. It is genuinely useful for comparing two designs assembled by the same method, and it is not evidence about how this equipment behaves in a container in west Texas.
Two consequences follow. First, a prediction is only as relevant as its reference conditions, and a figure computed at a benign reference temperature will overstate life in an enclosure whose ambient design range runs to +50 °C. Second, predictions cover the failure mechanisms in the model.
They do not cover installation workmanship, firmware defects, commissioning errors, connector fretting or a fleet-wide safety advisory — which on real projects account for a large share of the outage hours actually recorded. This page publishes no MTBF figures for BESS equipment for that reason: the numbers in circulation rarely carry the conditions, boundary and method that would make them comparable, and repeating one without them would be inventing precision.
Interrogating an MTBF in a technical proposal
Six questions settle most of it. What is the item boundary the number applies to. What counts as a failure. Is it calculated or field-observed, and if calculated, by which method and at what reference temperature. If observed, over how many device-hours and how many failures. Which convention — does it include repair time. And what duty is assumed, since a battery plant cycling twice a day loads contactors, fans and compressors at a rate an assumption of light duty will not reproduce.
Do not compare two vendors' MTBF figures until those six answers match. They almost never do, and an MTBF comparison across mismatched definitions is a comparison of drafting styles. Where the answers cannot be obtained, the honest position is that the figure supports nothing and the availability model should be built from the site's own event and restoration assumptions instead.
What MTBF is for on a battery project
It is a planning input, not a commercial term. It sizes spares holdings, sets expected annual event counts for the O&M budget, and feeds one half of the availability identity — with mean time to repair supplying the other and usually dominating the answer. Long-term service agreements guarantee availability, commonly 95-98 percent measured annually, along with retained capacity and round-trip efficiency. They do not guarantee MTBF, and a supplier who offers an MTBF figure in place of an availability commitment has offered you a statistic instead of a remedy.
One structural point carries over from failure rate and is worth restating here because it is so often inverted: for components in series, system MTBF is lower than any individual component MTBF, because the rates add and MTBF is their reciprocal. A plant assembled from parts each rated at hundreds of thousands of hours has a system figure far below any of them. That is not a defect in the equipment. It is what repetition does.
A 500,000-hour MTBF means the equipment should last about 57 years.
In reality: MTBF is a rate statistic for a population, not a service life for a unit. It says a fleet accumulating 500,000 operating hours expects roughly one failure, so 100 units over one year — 876,000 device-hours — expect between one and two. Under the constant-hazard assumption the number depends on, an individual unit's chance of reaching t = MTBF without a failure is exp(-1), about 37 percent. Service life is set by wear-out mechanisms that the flat-hazard model explicitly does not describe, which is why manufacturers publish design life separately when they publish it at all.
- Availability Glossary
- Failure rate Glossary
- Mean time to repair Glossary
- BESS procurement and contracts Article
Mean time between failures, in context.
The Grid-Scale BESS course covers mean time between failures — and the rest of the system — from the ground up, the way it actually gets deployed.