Commercial

Reliability run

A reliability run is the last commissioning test before a battery plant is accepted: continuous or near-continuous operation over a window of days to a few weeks, measured against an availability threshold written into the contract, with defined rules for which interruptions reset the clock.

It proves what no component test can — that the plant holds together as one machine over consecutive days, through diurnal thermal cycling, real dispatch instructions and whatever the site throws at it.

It measures readiness, not energy, so it is a different test from the capacity test that usually runs beside it. Passing it gates the commercial operation date. Its threshold and the operating-year availability guarantee are separate numbers proven by separate procedures, and the three definitions that decide the outcome — what counts as an interruption, what is excluded, and what resets the clock — are contractual, not statutory.

Reviewed August 2026 by Sergey Syrvachev

New to BESS? Start free with the 7-email fundamentals course — no cost, no account.

What the run demonstrates that nothing else does

Every test before it is bounded. The factory test exercises equipment against its specification, the site acceptance test checks the installation against codes and installation instructions, the capacity test integrates one discharge, the P-Q test walks a matrix of operating points. None of them runs the plant for a week.

The reliability run does, and what it catches is the class of fault that only appears with time in it: a thermal loop that trips on the fourth consecutive hot afternoon, a communications path that drops once a day, a firmware watchdog that reboots a converter on a schedule nobody documented, a spares process that takes eleven hours to get a technician through the gate with the right board.

The metric is availability over the window, computed the way the contract says — normally available hours over period hours less excluded hours, the same shape the operating-year guarantee uses. What differs is everything around the formula. The window is short, the plant is new, and the exclusions list is often not the annex negotiated for the operating years. Two documents can both say 97 percent and impose materially different obligations.

What counts as an interruption

The definition has to answer three questions, and money moves on each. What capability was lost: a full trip is unambiguous, a partial derate is not, so the run needs the same threshold rule availability needs — above some fraction of guaranteed capability a derated plant counts as available, below it as unavailable — or capacity-weighted proration that scores the derate pro rata. One power block out of ten, under a time-based metric with a permissive threshold, can score near 100 percent while the plant is visibly short.

Whose fault: grid outage, curtailment instruction, owner-caused delay and force majeure are the usual carve-outs, and a broad phrase about events beyond supplier control, with no enumerated list behind it, will be read generously by the party that drafted it. And how long: a protection trip that auto-recloses in forty seconds and a container down for a day are both interruptions, but a run scored on one-minute SCADA resolution and one scored on hourly averages will not agree about the first.

The fourth question is whether the plant was ever asked to do anything. A window in which no dispatch instruction arrived proves the plant sat there without tripping, which is standby rather than response. A defensible procedure requires exercised charge and discharge, a minimum number of full-power transitions in both directions, and a state-of-charge readiness rule — an asset parked full is perfectly available on paper and physically unable to absorb when the instruction says charge.

The same percentage is a far tighter obligation over two weeks — and the run scores a brand-new plant during infant mortality, on a different exclusions list.
97% over a 14-day reliability run336 period hours — 1 point ≈ 3.4 h~10 hours, total97% over an operating year~263 hours10 h100 h263 hunexcused downtime allowed at 97% availability

Three definitions decide the run before any hardware does: what counts as an interruption, which events are excluded, and which reset the clock rather than pause it. A time-based metric with a permissive partial-capability threshold can read near 100% with a power block down.

Key facts
What it is
A sustained-operation demonstration over days to a few weeks against a contractual availability threshold; passing it gates COD
Metric shape
Available hours ÷ (period hours − excluded hours) — the same formula as the operating-year metric, on a different window with a different exclusions list and a new plant
Window arithmetic
A 14-day run is 336 period hours: 1 point of availability ≈ 3.4 h, and a 97% threshold allows ~10 h of unexcused downtime for the whole window
Against the annual number
97% over a year ≈ 263 h, 98% ≈ 175 h, 95% ≈ 438 h — the same percentage is a far tighter obligation over two weeks
Three definitions decide it
What counts as an interruption, which events are excluded, and which reset the clock rather than pause it
Partial capability needs a rule
A threshold above which a derated plant counts as available, or capacity-weighted proration — a time-based metric with a permissive threshold can read near 100% with a power block down
Dispatch requirement
A window with no instructions proves standby, not response — require exercised charge and discharge and an SOC-readiness rule
Reference capability
The run must name the capability it is scored against; on a phased site, per block, matching partial-COD mechanics
No governing standard
Availability vocabulary borrows from IEEE Std 762 and is adapted contract by contract — the run's threshold, exclusions and reset rules come from the contract alone
Status of the energy
Under the FERC pro forma LGIA, pre-COD on-site test operations are Trial Operation, and electricity generated then is excluded from Commercial Operation

Reset, pause, or count

Three treatments exist for any event, and the contract has to assign one to every category. It can pause the clock and resume where it stopped, which is what excluded events normally do. It can count against the threshold and let the run continue, which is what ordinary faults do inside the allowance. Or it can restart the run from zero, the sanction reserved for events the owner will not accept at all — a safety-system actuation, a fire-alarm activation, a trip traced to a defect in a scope the supplier owns.

A loosely written reset trigger is what puts COD out of reach. On a new plant working through infant mortality, a broadly written reset trigger can loop the run indefinitely while the guaranteed COD passes and delay damages accrue daily — which is exactly when both parties stop fixing the plant and start litigating the classification of each event.

The durable drafting is what works for the annual metric too: an enumerated closed list of what resets, one named data source at a stated resolution, and a cap on how many restarts the parties will tolerate before they have to agree on something else.

Why the commissioning threshold is not the annual guarantee

Convert both to hours and the gap is obvious. A 14-day run is 336 period hours, so one point of availability is about 3.4 hours and a 97 percent threshold allows roughly 10 hours of unexcused downtime for the entire window. The operating-year guarantee at the same 97 percent allows about 263 hours; 98 percent allows about 175 and 95 percent about 438. A single eight-hour repair is a rounding error against the annual figure and most of the whole allowance against the two-week one. Short windows punish clustered events, and a commissioning window is the most clustered period in the asset's life.

The populations differ as much as the arithmetic. The run scores a plant in its infant-mortality period, with punch-list work still open and a crew still learning it; the guarantee scores a mature plant over a year, against a negotiated exclusions annex, frequently with a reduced or ramped year-one number precisely because holding a brand-new site to steady-state figures produces a dispute rather than a well-run plant. Carrying an argument from one into the other is a category error. Passing the run is evidence the plant works. It is not evidence the annual guarantee will hold.

Running it on a plant that is not finished

The most common dispute in this test is the denominator. If a block is not yet energised, if the collection system is only partly available, if the ISO has not opened a market-qualification slot, or if the site is still on a temporary auxiliary feed, then available has to be measured against something — full contracted capability, or the capability that physically exists. Say which, in writing, before the window opens. On a phased site the answer is usually per block, matching the partial-COD mechanics the offtake already needs, with the threshold and the reset rules applied block by block.

The second dispute is who may touch the plant. If commissioning work continues through the run, the supplier's own crew generates the interruptions the supplier is being measured on, and each one becomes a classification argument. Freeze the configuration, log every intervention with its authorisation, and decide in advance whether a firmware push during the window is a reset event.

Note where the run sits legally, too: under the FERC pro forma Large Generator Interconnection Agreement, on-site test operations and commissioning before Commercial Operation are Trial Operation, and electricity generated during Trial Operation is expressly excluded from commercial output — so the run is executed on commissioning energy, under whatever arrangement the interconnection agreement and the offtake made for it.

Common misconception

The reliability run is the availability guarantee tested early — pass it and the annual number is proven.

In reality: They are separate numbers proven by separate procedures. The run scores a brand-new plant over days to a few weeks, during infant mortality and often with punch-list work still open; the guarantee scores a mature plant over a year against a negotiated exclusions annex, frequently with a reduced or ramped year-one figure. The arithmetic alone separates them: 97 percent over 14 days allows about 10 hours of downtime, while 97 percent over a year allows about 263. Passing the run proves the plant holds together as one machine; it says little about whether the annual guarantee will hold.

Go deeper

Reliability run, in context.

The Grid-Scale BESS course covers reliability run — and the rest of the system — from the ground up, the way it actually gets deployed.

Browse the course