Performance

Failure rate

Failure rate is λ, failures per unit of operating time. It is quoted per hour, per year, or in FIT — one FIT being one failure per 10^9 device-hours, a unit that exists because 10^9 hours is about 114,000 years and the per-hour figure otherwise arrives as a string of zeros. λ is not a constant of nature for a piece of equipment.

It is high early while manufacturing and installation defects burn off, roughly flat through useful life, and rising again at wear-out — the bathtub. Only in that flat middle does λ = 1/MTBF hold, and only there does R(t) = exp(-λt) describe survival. The property that matters most on a battery plant is that λ adds: put components in series and the rates sum, so system MTBF is lower than any single component's.

Reviewed August 2026 by Sergey Syrvachev

New to BESS? Start free with the 7-email fundamentals course — no cost, no account.

Units, and what a FIT figure is actually counting

A rate of 2.3 × 10^-7 failures per hour and 230 FIT are the same number written twice. FIT scales the quantity to a billion device-hours so that semiconductor and component data stay readable, and it inverts straight back: 230 FIT is an MTBF of about 4.3 million hours for that item.

The inverter-topology entry on this site already works in FIT, citing Semikron Danfoss figures for cosmic-ray-induced semiconductor failures — around 230 FIT per switch for a two-level 1200 V module blocking 1000 V, against below the roughly 1 FIT measurement limit for a three-level part blocking 500 V, with a TNPC design landing near a third of the two-level rate.

Read that example for what it is. Those figures cover one failure mechanism, cosmic-ray-induced single-event burnout, at one blocking voltage. They are not the failure rate of a PCS, or of a phase leg, or of anything you could put in an availability model on their own. A component λ is always λ for a stated mechanism under stated conditions, and adding two rates measured under different conditions produces a number that means nothing. The conditions travel with the rate or the rate is not usable.

The bathtub curve is three different physical stories

The decreasing-hazard region is populated by things that were wrong from the start: a marginal weld in a cell-contacting system, a torque that was never checked, a sense line landed on the wrong channel, a firmware parameter left at a factory default. Factory and site acceptance testing exist partly to move these failures forward in time, and the availability entry's observation that year one usually runs below the steady-state guarantee is the same phenomenon written into a contract.

The flat middle is where the exponential assumption earns its keep, and where λ = 1/MTBF is a legitimate shorthand. The rising end is wear-out, and on a BESS it is largely mechanical and electrolytic rather than electrochemical: HVAC compressors and pumps, fans, contactors with a finite operation count, DC-link capacitors, breaker mechanisms. Those items have design lives, and a plant that has run twelve quiet years is not entitled to assume the next twelve look the same. Treating a single λ as valid across the whole life is the error that makes long-horizon availability models optimistic.

λ is not a constant of nature — it falls, flattens and rises, and the equations everyone uses are valid only in the flat middle.
ONLY HERE: λ = 1/MTBF and R(t) =exp(−λt)infant mortality —manufacturing andinstallation defectsburning offuseful lifewear-out — compressors,pumps, fans, contactorcounts, DC-link capacitors,breaker mechanismsfailure rate λno numeric scale: the body gives the SHAPE, not valuesUnits: per hour, per year, or FIT — and 1 FIT is one failure per 10⁹ device-hours, about 114,000years, which is why the unit exists at all.

Rates add in series — λ_sys = Σλi — so the plant’s aggregate event rate is set mainly by how many parts there are, and system MTBF is lower than any component MTBF. Good design changes what each event removes, by bounding fault domains, rather than how often events arrive. A worked conversion for scale: 230 FIT is 2.3 × 10⁻⁷ per hour, about 4.3 million hours MTBF for that one mechanism; six such switches sum to roughly 1,380 FIT, about 725,000 hours. That site-cited example is a two-level 1200 V module blocking 1000 V at about 230 FIT per switch for cosmic-ray failures, against a three-level part at 500 V sitting below the ~1 FIT limit (Semikron, one mechanism only).

Key facts
Symbol and units
λ, failures per unit time — per hour, per year, or in FIT
FIT definition
1 FIT = 1 failure per 10^9 device-hours; 10^9 hours ≈ 114,000 years
Conversion
230 FIT = 2.3 × 10^-7 /h ≈ 4.3 million hours MTBF for that mechanism
Site-cited example
2L 1200 V module blocking 1000 V ≈ 230 FIT/switch for cosmic-ray failures; 3L at 500 V below the ~1 FIT limit (Semikron, one mechanism only)
Series summation
λ_sys = Σλi, so system MTBF is lower than any component MTBF — six switches at 230 FIT ≈ 1380 FIT ≈ 725,000 h
Bathtub curve
Decreasing (infant mortality) → flat (useful life) → increasing (wear-out); λ constant only in the middle
When λ = 1/MTBF applies
Constant-hazard region only — the same assumption that makes R(t) = exp(-λt) valid
BESS wear-out items
HVAC compressors and pumps, fans, contactors with finite operation counts, DC-link capacitors, breaker mechanisms
Design consequence
With thousands of repeated parts the aggregate rate is set by count — bounded fault domains beat component-level heroics

Series summation, and why system MTBF is always the smallest number

For n independent components in series — meaning any one failure removes the function — the system rate is the sum: λ_sys = λ1 + λ2 + … + λn, and system MTBF is 1/λ_sys. That result is lower than the MTBF of every component in the chain, including the best one. It is not an average and it does not improve when you add a good component; adding anything in series adds rate.

Work it with the audited figure above. One switch at 230 FIT is about 4.3 million hours for that mechanism. A three-phase two-level bridge has six switches in series in the sense that any one of them failing takes the converter down, so the rate is 6 × 230 = 1380 FIT and the converter MTBF for that mechanism alone falls to roughly 725,000 hours. Six components, one-sixth the MTBF. Now scale it to a plant with hundreds of power stages and the arithmetic stops being abstract.

On a BESS, count is the dominant term

A grid-scale battery plant is built by repetition. Hundreds of thousands of cells, thousands of modules, hundreds of racks each with its own contactor and fuse, dozens to hundreds of power stages, an HVAC fleet, sense lines, comms nodes. Even with excellent components, the aggregate event rate is high by construction, and no procurement decision about a single part changes that. The site will generate field events every year and the operating plan has to assume it.

Which is why the design lever is not λ but the size of the fault domain. Rack-level isolation so a faulted string drops without de-energising a container. Many small power blocks rather than one large one, so a converter fault costs a slice — the inverter-topology entry records Sungrow's claim of an 8 percent output loss on a string PCS fault with 11 of 12 units still online, which is a vendor white-paper figure rather than a measurement, but the architecture argument behind it is sound.

Under a capacity-weighted availability metric those bounded slices are exactly what the number rewards, and the availability entry's worked case — one of ten power blocks down for a month costing about 0.83 points of annual energy availability — is the same idea priced.

Reading a λ someone hands you

Ask where it came from. A prediction built from a parts-count standard is a spreadsheet exercise against reference conditions; a field rate is observations divided by accumulated device-hours and carries a confidence interval that a point estimate hides.

Ask what the item boundary is — die, module, converter, container — because the same word covers all four in different documents. Ask what counts as a failure. Ask what temperature, duty and voltage the rate is stated at, since the FIT example above changes by more than two orders of magnitude between two blocking voltages of the same technology family.

Then ask what you are allowed to do with it. A per-mechanism component rate feeds a component comparison. It does not feed a plant availability model unless every other contributing mechanism has been counted at the same boundary under the same conditions, and on a real site the mechanisms that dominate outage hours are usually not the ones with published rates.

Common misconception

Specify high-reliability components and the plant inherits a low failure rate.

In reality: Rates add. For components in series the system rate is the sum of the component rates, so system MTBF is lower than the MTBF of every part in the chain, and it falls further with every item added. A battery plant is thousands of parts in a mixture of series and parallel arrangements, so its aggregate event rate is set mainly by how many things there are. What a good design changes is not the arrival rate of events but the capability each one removes — which is why rack-level isolation and many small power blocks do more for the availability number than an upgraded component ever will.

Visuals & further reading
Go deeper

Failure rate, in context.

The Grid-Scale BESS course covers failure rate — and the rest of the system — from the ground up, the way it actually gets deployed.

Browse the course