Performance

Unplanned outage

An unplanned outage is a loss of capability the plant did not schedule — a PCS trip, a BMS-commanded contactor open, an HVAC failure that forces a thermal derate, an MV breaker or protection operation, a comms loss that drops the site out of dispatch, or a grid-side event. The generating-unit vocabulary splits the unscheduled family in two.

A maintenance outage can be deferred to a convenient window but not to the next planned one; a forced outage cannot wait that long, and is graded by how urgently the unit had to come off — immediate, delayed, postponed, with a failure to start as its own class.

On a modular battery plant the more useful distinction is different. Most events remove a slice, not the site, and whether that slice shows up at all depends on whether the availability metric is time-based with a threshold or capacity-weighted. Same event, two very different numbers.

Reviewed August 2026 by Sergey Syrvachev

New to BESS? Start free with the 7-email fundamentals course — no cost, no account.

What actually takes capability off a battery plant

The PCS trips on DC or AC overcurrent, on a DC bus excursion outside the VDC window, on a grid disturbance beyond its ride-through envelope, on a stack or gate-drive fault, or on loss of its own auxiliary supply.

The BMS commands contactors open on a cell voltage or temperature limit, an isolation-resistance fault, or a lost sense line — and short of an open contactor it imposes current derates, which the battery-management-system entry identifies as the most common reason a plant delivers less than nameplate on a summer afternoon. A derate is not an outage, and where it lands in the availability calculation is a drafting question rather than an engineering one.

Thermal faults escalate on a slower clock than electrical ones. A tripped chiller loop derates a container within hours and eventually forces a shutdown, because continuing to cycle outside the warranted temperature envelope trades an availability problem for a warranty problem. Then the electrical plant: an MV feeder fault, a protection relay operating correctly on a real fault or incorrectly on a mis-set curve, a switchgear mechanism failure, loss of the auxiliary transformer that feeds HVAC and controls.

And comms — a SCADA or market-link failure leaves a physically healthy plant that cannot be dispatched, with the grid code or market rules deciding whether it holds last setpoint, ramps to zero, or follows a local default. Whether that state is an outage has to be written down; it is one of the most common definitional gaps in an availability annex. Grid-side events — a POI or gen-tie outage, a curtailment instruction, utility work — are usually excluded, but only if the exclusions list enumerates them.

Most events are partial, and that is the whole measurement fight

A grid-scale BESS is built as many bounded fault domains, so the typical event removes one rack, one container or one power block. A time-based metric with a threshold rule can score that at or near 100 percent while real capability is missing; a capacity-weighted metric prorates it. The availability entry's worked case makes the gap concrete: one of ten power blocks down for a month is about 0.83 points of annual energy availability, and close to nothing under a thresholded time-based definition. Run twelve such events in a year and the two metrics describe different plants.

Architecture sets how big the slice is. The inverter-topology entry records Sungrow's claim of an 8 percent output loss on a string PCS fault because the remaining 11 units in the container stay online — a vendor white-paper figure rather than a witnessed measurement, but the structural point holds against a central converter whose fault takes its block to zero.

Which is why the availability formula has to be chosen to match the single-line diagram. Capacity-weighted proration with an enumerated, closed exclusions list and one named data source at a stated resolution is the durable middle ground the availability entry argues for, and it is on partial events that the argument is won.

On a modular plant most unplanned events remove a slice, not the site — and whether the metric sees slices is a contract clause, not a physical fact.
capacity-weighted metricone of ten power blocks, one month~0.83 pointstime-based, with a thresholdthe plant lost the same energyreads ~zero0.51.0 pointannual availability lost to one block down for a month

The raw forced-outage rate misleads here too: a BESS sits in reserve shutdown most of the year, so few service hours shrink the denominator and inflate the rate while saying nothing about readiness in the hours that mattered. The other hard case for the annex is a healthy plant that cannot be dispatched because communications are down.

Key facts
Definition
Unscheduled loss of capability from a failure, trip, or protective action
IEEE Std 762 heritage
Three outage classes: planned, maintenance (deferrable, but not to the next planned window) and forced — the last graded immediate, delayed, postponed, plus failure to start; adapted contract by contract
Forced outage rate
Forced outage hours ÷ (forced outage hours + service hours); equivalent forms convert derated hours to equivalent full-outage hours
Why the raw rate misleads
A BESS sits in reserve shutdown most of the year, so few service hours shrink the denominator and inflate the raw rate — which still says nothing about readiness in the hours that mattered; use the equivalent and demand-weighted forms
Typical causes
PCS trip, BMS contactor open or current derate, HVAC failure, MV protection operation, auxiliary-power loss, comms loss, grid-side events
Partial by construction
1 of 10 power blocks down for a month ≈ 0.83 points of annual energy availability — near zero under a thresholded time-based metric
Fault-domain claim
Sungrow claims 8 percent output loss on a string PCS fault with 11 of 12 units online — a vendor figure, not a measurement
Definitional gap
A healthy plant that cannot be dispatched because comms are down — outage or not? The annex must say
Biggest restoration lever
Whether a fault latches out and needs a local reset, or clears remotely
Attribution
Root cause decides whether the LTSA, the O&M agreement, an exclusion, or the owner carries the hours

Forced-outage vocabulary, and why the raw version misleads here

There is no BESS-specific availability standard, so the terminology is borrowed from IEEE Std 762, the generating-unit reliability vocabulary, and adapted contract by contract. Its core rate is the forced outage rate: forced outage hours divided by the sum of forced outage hours and service hours. The equivalent forms extend it by converting derated hours into equivalent full-outage hours — which is precisely the extension a battery plant needs, because most of its events are derates rather than removals.

The raw rate has a structural problem on this asset class, and it runs the opposite way from intuition. Reserve-shutdown hours fall outside both halves of the ratio, so a resource that spends most of the year not called accumulates very few service hours and posts an inflated rate for the same outage hours: 100 forced outage hours against 8,000 service hours is 1.2 percent, the same 100 hours against 300 service hours is 25 percent. A BESS is that resource by construction — it sits idle for most of 8,760 hours and earns in the few hundred that count.

What the raw rate never tells you is whether the plant was ready in the hours that were worth money, and that is what the demand-weighted variants in the same vocabulary are for: they weight forced outage hours by the probability the unit was actually needed, so on a reserve-shutdown-heavy asset they land below the raw rate rather than above it. Importing a thermal-fleet forced outage rate into a storage contract without the equivalent and demand-weighted treatment measures the wrong thing confidently.

The first hour decides most of the down time

Detection and diagnosis dominate on an unattended site, and the sequence-of-events file and historian are both the diagnostic tool and, later, the evidence — the SCADA entry makes the point that the historian is the record every availability argument gets fought with. Time synchronisation across hundreds of devices is what makes a sequence of events readable at all; without it, a cascade of trips arrives as an unordered list and the root cause is a guess.

The single highest-leverage question on a modular plant is what requires a truck. A PCS that rides through, latches, and can be reset remotely turns a grid disturbance into a blip. The same fault on a unit that requires a local reset at a rural site turns it into a day, and if a dozen units latch together, into a week of block-by-block recovery. Ask, in writing and before procurement closes: which fault codes are remotely resettable, which require a site visit, and which require the OEM. That answer moves restoration time further than any component specification.

Who pays depends on the root cause

Attribution is where the money lands. The same lost hours are an LTSA guarantee breach if the battery system failed, an O&M matter if the balance of plant failed, an excluded event if the grid caused it, and the owner's own problem if dispatch emptied the plant. The availability entry makes this point about a BMS-imposed derate specifically: the root cause decides whether the lost megawatts are a breach, an exclusion, or another contract's problem. Which is why root-cause analysis on a significant event is a commercial process with an engineering method, not the other way round.

Two patterns deserve their own clause. Systemic events, where one firmware defect or one bad component batch produces sixty simultaneous outages — most contracts do not say whether that is one event or sixty, and cap language behaves very differently under each reading. And year one, where infant-mortality faults and commissioning punch-list work run the rate above steady state; many contracts run a reduced first-year guarantee or a defined ramp period for that reason, which is better than holding a new site to a mature number and guaranteeing a dispute instead of a well-run plant.

Common misconception

An unplanned outage means the site went down.

In reality: On a modular battery plant most unplanned events remove a slice — one rack, one container, one power block — while the rest of the site keeps operating. Whether that slice appears in the availability number depends entirely on the metric: a time-based definition with a partial-capability threshold can report near 100 percent through a year of such events, while a capacity-weighted definition prorates every one of them. The plant lost real energy in both cases. Only one of the two measurements noticed.

Visuals & further reading
Go deeper

Unplanned outage, in context.

The Grid-Scale BESS course covers unplanned outage — and the rest of the system — from the ground up, the way it actually gets deployed.

Browse the course