Performance

Mean time to repair MTTR

MTTR is the average time taken to restore a failed item to service, and in the steady-state identity A ≈ MTBF / (MTBF + MTTR) it is the term an owner can actually move. The trap is what the clock contains. A restoration spans detection, diagnosis, crew mobilisation, waiting for the part, the physical repair, and return to service including any re-test — and most contractual definitions of MTTR cover only the hands-on work.

Logistics delay is commonly carved out and tracked separately as mean logistics delay time. That single drafting choice is how a supplier reports an excellent MTTR on a plant that sat dark for six weeks, and it is why the number to argue about is usually mean down time, not MTTR.

Reviewed August 2026 by Sergey Syrvachev

New to BESS? Start free with the 7-email fundamentals course — no cost, no account.

Break the clock into segments before you argue about the number

Detection comes first, and on an unattended rural site it is not free. The alarm has to reach a human, be recognised as real rather than one of the day's nuisance annunciations, and be timestamped in a historian someone trusts. A fault that self-clears and re-arms can hide for days while capability quietly stays down.

Diagnosis follows: fault code to root cause, done remotely if the telemetry supports it and by a site visit if it does not. Mobilisation is travel time to the site, plus permits, switching orders, lock-out/tag-out, and the availability of people qualified for the voltage class involved.

Then logistics delay — is the part on site, at a regional depot, or at a factory? — then the repair itself, then return to service. That last segment is routinely underestimated: re-energisation sequencing, a functional re-test, updating the historian, and where a market requires it, re-registration or a re-qualification run before the resource can be dispatched again. Six segments, and on most BESS events the hands-on repair is not the longest of them.

Logistics delay is usually outside MTTR, and that is the whole point

Reliability practice separates active repair time from logistic delay time and administrative delay time deliberately, because they are owned by different people and fixed by different means. A supplier controls its technicians' efficiency; it does not control whether the owner funded a spare. So a contract that guarantees MTTR alone has guaranteed the segment the supplier controls and left the segments that dominate the outage unmeasured.

The consequence is direct: a plant that fails rarely but waits six weeks for a replacement transformer posts worse availability than one that fails more often and recovers in hours.

Put the site's own lead-time figures beside that — roughly 12 to 18 months for MV units and 24 to 36 months or longer for HV station transformers — and it is clear that for the large single-point items, restoration time is a procurement decision made years before the failure. When you want a number that ties to availability, ask for mean down time with an enumerated segment list, and treat MTTR as one line inside it.

An excellent MTTR is entirely compatible with terrible availability — because the segment the months live in is the one commonly carved out of it.
detectionthe clock startsherediagnosismobilisationlogistics delay12–18 mo for MV tx24–36 mo for HV txphysical repairreturn to serviceincluding re-testcontractual MTTRcommonly carved outcontractual MTTRmean down time — what the owner actually waits

MTTR is the movable term in A ≈ MTBF / (MTBF + MTTR): the failure rate is largely fixed by the equipment, and the restoration time is what an owner can actually buy. Ask for mean DOWN time instead, with an enumerated segment list and a named clock start and stop, and read the carve-outs before the headline number. Restoration is mostly a spares decision made years earlier — HV station transformers have run 24–36 months or longer against roughly 12–18 for MV units — and repair granularity is fixed at design: a module swap where the cell-to-cell connection is welded in, rack isolation by contactor, a PCS stack swap. None of those is a choice available on the day.

Key facts
Role in the identity
The movable term in A ≈ MTBF / (MTBF + MTTR)
Segments of a restoration
Detection, diagnosis, mobilisation, logistics delay, physical repair, return to service including re-test
The carve-out
Logistics delay is commonly excluded from contractual MTTR and tracked as mean logistics delay time
Consequence
An excellent MTTR is compatible with terrible availability — a six-week wait for a replacement transformer
Ask for instead
Mean down time, with an enumerated segment list and a named clock start and stop
Long-lead exposure
MV transformers ~12-18 months, HV station transformers 24-36 months or longer — restoration is a spares decision made years earlier
Repair granularity
Module swap (CCS is welded in), rack isolation by contactor, PCS stack swap — all fixed at design
Contract lever
Guaranteed response times bind; MTTR is reported after the fact
Affordability check
97 percent ≈ 263 h/yr; across a dozen full-site events that is ≈ 22 h mean down time each
Scope seam
LTSA holds battery spares, OMA holds balance-of-plant spares — write the supply and response duties across the seam

What actually sets restoration time on a battery plant

Spares, and the granularity of the repair unit. Both were decided at design and procurement, not at the fault. Module-level replacement is the common repair for a battery failure, because the cell-contacting system is welded into the module — a field failure means pulling and swapping the module, not repairing a connector.

Rack-level contactors let a faulted string be dropped without de-energising a container, which converts a would-be outage into a bounded derate. A PCS with a replaceable power stack is restored in a shift; one that needs a factory return is not. Fast-moving items belong on site; the critical spares are the ones whose absence strands a whole block, and the transformer case is the extreme.

The seam between contracts matters as much as the shelf. The LTSA typically covers the battery system and the O&M agreement covers the balance of plant, so the party holding the spare and the party owing the availability guarantee are not automatically the same party. If the OMA holds the MV spares and the LTSA carries the availability number, then either the response and supply obligations are written across the seam or the guarantee is unenforceable in exactly the cases that cost most.

Writing restoration into a contract so it binds

Guaranteed response times are enforceable; MTTR is a statistic reported after the event. Contract the first and measure the second. A workable clause names the clock start — which alarm, in whose historian, at what resolution — and the clock stop, distinguishing back-in-service from back-in-service-after-re-test. It enumerates what is excluded rather than gesturing at events beyond supplier control: a safety hold after a thermal event, an AHJ re-inspection, a utility outage window the owner cannot schedule, weather that prevents crane work.

The three letters are also read three ways in the field — repair, restore, respond — and occasionally a fourth, resolve. They mean different things and the difference is hours to weeks. Define the term in the agreement rather than assuming the industry has a shared meaning, because it does not.

How much restoration time the guarantee can afford

Run the arithmetic backwards from the availability figure. Convert the common guarantee band into hours first: 98 percent allows roughly 175 hours of unexcused downtime a year, 97 percent about 263 hours, 95 percent about 438. Assume — and it is your assumption to make, not a published statistic — that a site sees a dozen unplanned events in a year, each taking the whole plant off.

At 97 percent, mean down time per event must sit under about 22 hours; at 95 percent, under about 36. Change the event count and the answer moves, which is the point: the two terms are traded against each other, and a supplier promising fewer events is promising something you cannot verify inside a single operating year.

Modularity changes the sum in the plant's favour, and this is where a capacity-weighted metric earns its place. If an event removes one of ten power blocks rather than the site, the same wall-clock restoration costs a tenth of the availability. At the far end of that arithmetic, one block of ten down for a month costs about 0.83 points of annual energy availability. Restoration speed and fault-domain size multiply; improving either one improves the number, and improving both is what separates a plant that hits 97 percent from one that argues about exclusions.

Common misconception

MTTR is how long the plant is down after a failure.

In reality: MTTR usually measures only the hands-on restoration. The plant's actual down time also contains detection, diagnosis, crew mobilisation, waiting for the spare, and return to service with any re-test, and most contracts carve logistics delay out into a separate mean logistics delay time. That is exactly how a supplier reports a few hours of MTTR on an event that kept a block offline for weeks. If the number you want is the one that drives availability, ask for mean down time with every segment enumerated.

Go deeper

Mean time to repair, in context.

The Grid-Scale BESS course covers mean time to repair — and the rest of the system — from the ground up, the way it actually gets deployed.

Browse the course