Reliability
Reliability is the probability that an item performs its required function without failure over a stated period, under stated conditions. The two statements are not decoration. A converter described as 99 percent reliable means nothing until you know 99 percent over what — a year, twenty years, one dispatch — at what ambient temperature, under what duty, and against what definition of failure.
Where the hazard rate is constant, the probability takes the exponential form R(t) = exp(-λt), and it is that assumption which makes the exponential valid, not the other way round. Reliability is also not availability. Reliability asks whether an item fails; availability asks whether the plant is ready when the instruction arrives. On a BESS the contract stack guarantees the second and almost never the first.
Reviewed August 2026 by Sergey Syrvachev
New to BESS? Start free with the 7-email fundamentals course — no cost, no account.
The period and the conditions are half the number
Quote a reliability figure without a period and you have quoted nothing, because R falls monotonically with time. The same item can be 0.999 over a month and 0.988 over a year, and both are true statements about one λ (0.999 raised to the twelfth is 0.988).
Quote it without conditions and you have quoted something worse than nothing, because the conditions are what λ was measured or predicted under. A cabinet fan rated in a 25 °C laboratory is not the same fan inside an enclosure holding cells near 25 °C while rejecting heat to a 45 °C summer ambient, and the container thermal-management entry sets out how far those two environments diverge in practice.
The third quiet input is the definition of failure. Does a PCS that rides through a grid disturbance, latches out and needs a manual reset count as a failure? Does a module the BMS has isolated while the rack keeps running? Does a derate? On a modular battery plant a large share of events remove capability without removing the plant, so a reliability statistic assembled under a strict failure definition and one assembled under a loose one describe the same hardware with numbers that differ by an order of magnitude. Ask for the definition before you compare two vendors' figures.
Where the exponential form comes from, and where it stops holding
R(t) = exp(-λt) is not a general law of equipment. It is the solution you get when the hazard rate is constant — the assumption that the item is as likely to fail in its next hour of service as it was in its first. Only under that assumption does λ = 1/MTBF hold, and only under it can a single number stand in for a whole life.
Real equipment follows a bathtub hazard curve: a decreasing region where manufacturing and installation defects burn off, a roughly flat middle, and a rising wear-out region as bearings, contactors, capacitors and compressors reach the end of their design lives. Infant mortality and wear-out both violate the constant-hazard assumption, which is exactly why the exponential works well in the middle of life and badly at both ends.
One consequence of the exponential form is worth carrying around, because it corrects a common instinct. Survival to t = MTBF is exp(-1), about 37 percent — not 50 percent, and not a near-certainty. If the flat-hazard assumption applies at all, roughly two-thirds of a population will have failed at least once by the time the fleet has accumulated one MTBF each. That is the arithmetic that turns a large-sounding hours figure into an ordinary expectation of field events.
The definition carries two clauses and both are part of the claim: the probability of performing without failure over a stated period, under stated conditions. The two measures are bridged by A ≈ MTBF / (MTBF + MTTR). Reliability composes in series — R_sys = R1 × R2 × … × Rn — so adding a series item can only lower system reliability. What contracts actually guarantee is availability, commonly 95–98% annually, which is not a reliability figure. The one place a reliability TEST appears is the commissioning reliability run that gates commercial operation, and that is separate from the operating-year guarantee.
- Definition
- Probability of performing without failure over a stated period under stated conditions — both statements are part of the claim
- Exponential form
- R(t) = exp(-λt), valid only where the hazard rate is constant (the useful-life region)
- Survival at t = MTBF
- exp(-1) ≈ 37 percent, not 50 percent — most of a population has failed once by then
- λ = 1/MTBF
- Holds under constant hazard only; infant mortality and wear-out both break it
- Series arithmetic
- R_sys = R1 × R2 × … × Rn — adding a series item can only lower system reliability
- Reliability vs availability
- No-failure probability vs ready-when-asked; bridged by A ≈ MTBF / (MTBF + MTTR)
- What contracts guarantee
- Availability, commonly 95-98 percent annually — not a reliability figure
- Where a reliability test appears
- The commissioning reliability run that gates COD, separate from the operating-year guarantee
- Failure definition
- Whether a latched trip, an isolated module or a derate counts as a failure changes the statistic by an order of magnitude
Reliability and availability pull apart in both directions
The two metrics answer different questions and can disagree sharply. A plant that fails often but recovers in an hour — a modular site with rack-level isolation, remote reset and spares on the shelf — can post excellent availability on a poor reliability record. A plant that fails almost never can post terrible availability if the one failure it does have takes six weeks to fix.
The availability entry makes exactly that point with a replacement transformer, and this site's transformer entries put the underlying lead times at roughly 12 to 18 months for MV units and 24 to 36 months or longer for HV station transformers. Nothing about component reliability rescues that.
The bridge between them is the steady-state identity the availability entry states: A ≈ MTBF / (MTBF + MTTR) for a repairable system. Reliability lives in the first term, restoration in the second, and the plant's number is the ratio. Which is why a spares strategy, remote diagnostics and a bounded repair unit usually move availability further than any specification you can write about a single component.
Why the contract stack buys availability and not reliability
Look through a BESS contract set and you will not find a guaranteed reliability figure. You will find performance guarantees in three families — retained capacity, round-trip efficiency, availability — with the availability guarantee at utility scale commonly landing at 95-98 percent measured annually, and mature LFP sites targeting 97 percent and above.
The reason is that availability is the thing the offtaker actually buys and the thing an EMS or SCADA historian can measure hour by hour, while reliability is a probability statement that cannot be settled from twelve months of one plant's data.
The one place a reliability-shaped test does appear is commissioning. Most contracts end with a reliability run: continuous or near-continuous operation over days to a few weeks against an availability threshold, with rules for which interruptions reset the clock. Passing it gates the commercial operation date. It is a separate number proven by a separate procedure from the operating-year guarantee, and carrying one into an argument about the other is a category error.
What reliability is actually for on a battery project
Reliability earns its place as a design input rather than a contract term. Components in series multiply: for independent items where any one failure takes the function down, R_sys = R1 × R2 × … × Rn, so system reliability is always lower than the worst element in the chain and never higher. Add a series item and you have subtracted reliability. That is the arithmetic behind every architecture decision on a battery plant — how many racks share a contactor, how many blocks sit behind one PCS, whether the main power transformer has anything behind it at all.
So the useful design question is not how to raise R but how to shrink the consequence of the events R predicts. Parallel redundancy, rack-level isolation, many small power blocks instead of one large one, and an auxiliary supply that does not take the site down with it. A plant built from thousands of repeated parts will generate events; the engineering choice is whether each event costs a slice or the site.
Reliability and availability are two ways of saying the same thing, so a highly reliable BESS is a highly available one.
In reality: They diverge in both directions, and on a battery plant they routinely do. A modular site with rack-level isolation, remote reset and spares on site can fail often and still post strong availability because every event is cleared in hours. A site that almost never fails can post poor availability when the one thing that failed is a transformer with a replacement lead time measured in months. Reliability describes how often failures arrive; availability is what you get after restoration time is folded in, roughly MTBF / (MTBF + MTTR).
- Availability Glossary
- Failure rate Glossary
- Mean time to repair Glossary
- BESS commissioning and capacity testing Article
Reliability, in context.
The Grid-Scale BESS course covers reliability — and the rest of the system — from the ground up, the way it actually gets deployed.