Commercial

Root cause analysis RCA

Root cause analysis is the structured investigation that asks why a component failed rather than which component failed. The difference is clearest on a real chain. A battery module derates and then trips on high cell temperature. The proximate cause is a failed fan in that container's cooling loop.

The root cause is the filter that sat two quarters past its service interval and loaded the motor, or a control setpoint that let the block run at a combination of duty and ambient the cooling was never sized for. Replace the fan and the plant reproduces the fault on the same schedule.

On this asset class the investigation carries commercial weight as well, because its conclusion sorts the event into warranty defect, excluded event, operator error or design fault — four verdicts with four different payers.

Reviewed August 2026 by Sergey Syrvachev

New to BESS? Start free with the 7-email fundamentals course — no cost, no account.

Proximate cause, root cause, and the layers between

Work the temperature trip down layer by layer and the ownership changes at almost every step. The module tripped because the BMS enforced its limit, which is the protection working. The cells got hot because heat rejection fell. Heat rejection fell because a circulation fan stopped. The fan stopped because its motor was drawing above rating against a loaded filter, or because the bearing reached the end of its life at the hours it was always going to reach it.

The filter was loaded because the interval came from a manual written for a benign site and was applied unchanged to a dusty one, or because the visit was deferred. Each of those sentences has a different owner: the OEM, the O&M contractor, the asset manager who approved the deferral, the designer who picked the interval.

The chain also branches, and taking the first branch on offer is how investigations go wrong. In the same event the fan may have been healthy and the block simply run harder and hotter than the cooling was specified for — which makes it a dispatch and setpoint question rather than a maintenance one, and moves it from the O&M ledger to the design or operating-envelope ledger. Test both branches against the data before committing to one. A cooling loop that has been marginal on hot afternoons all summer leaves a trace in the delta-T record whether or not a fan failed on the day.

Stop when the next why leaves the boundary of what the project can change. An analysis that terminates at the ambient was hot has stopped a layer early: ambient is an input, and the actionable cause is the specification, the setpoint or the interval that assumed a different one. Method matters less than people think — five whys, a fault tree, a cause-and-effect diagram all get you down the same ladder — and the discipline that actually separates a useful RCA from a filed one is whether every step is supported by a record rather than by a recollection.

Evidence has a shelf life

The first action on a significant event is a data freeze, before anything is reset, restarted or cleared. The reason is resolution, and the plant carries several different ones. The historian the availability calculation runs on is commonly around one to five minutes per the contract-defined methodology — fine for downtime accounting and completely blind to a converter event that lasted 200 milliseconds.

The records that resolve that event live in the PCS fault buffer, the protection relay's disturbance recorder and the BMS's high-rate log, all of which are ring buffers of finite depth. They roll. A plant restarted quickly to recover revenue can overwrite the only description of what happened, and no amount of later analysis puts it back.

Pull the whole set, not the obvious one: converter fault buffers and event logs, relay disturbance records, per-cell BMS voltage and temperature history around the event, HVAC and coolant trends for the preceding days, the SCADA event and alarm sequence with its timestamps, and the dispatch record showing what the plant was being asked to do.

Timestamp alignment is the recurring practical failure — devices on different clock sources put the same event in a different order, and a sequence-of-events reconstruction built on drifting clocks produces confident nonsense about which thing caused which.

Physical evidence has its own rules. Where the contract provides for failure analysis, the removed unit goes back intact and tagged rather than into a skip, and the site keeps photographs of the as-found condition before anything is disturbed. There is a genuine tension here between restoring revenue and preserving the record, and it should be resolved in the O&M procedures in advance — a written rule about what gets captured before a restart, with a named person able to authorise the restart, beats the argument that otherwise happens at 3 a.m. with money on the clock.

The report must also return one of four verdicts — warranty defect, excluded event, operator error or design fault — identical downtime, entirely different money.
the ROOTa filter two quarters past itsinterval, or a setpoint that letthe block run a duty the coolingwas never sized forthe PROXIMATE causea failed cooling fanwhat you seea trip on high cell temperatureReplace the fan and the plant reproduces the fault on the same schedule. "The ambient was hot"stops one layer early — ambient is an input; the specification is the cause.

Freeze the data before any reset: PCS fault buffers, relay disturbance records, per-cell BMS history and SCADA event sequences live in finite ring buffers that roll on restart. An RCA whose only artefact is a report has corrected nothing — it must change an interval, a threshold, the spares list or the operating envelope.

Key facts
What it finds
The underlying cause, not the failed component — the fan is the proximate cause; the overdue filter, or the setpoint that let the block run hot, is the root
Stopping rule
Stop at the deepest cause the project can change. A chain that ends at "the ambient was hot" stopped one layer early — ambient is an input, the specification is the cause
First action
Freeze the data before any reset or restart: PCS fault buffers, relay disturbance records, per-cell BMS history, HVAC trends, SCADA event sequence and the dispatch record
Resolution mismatch
The availability historian at ~1-5 min cannot reconstruct a converter event lasting 200 ms; those records live in finite ring buffers that roll on restart
Timestamps
Devices on different clock sources reorder the sequence of events — a reconstruction built on unsynchronised clocks is confidently wrong about causation
Four verdicts, four payers
Warranty defect, excluded event, operator error, design fault — identical downtime, entirely different money
Contract hooks
Named investigator and deadline, evidence-preservation obligation, owner or IE attendance and raw-data rights, availability treatment while open, and an escalation route for conflicting conclusions
What it must change
An interval (PM), a threshold (PdM), the spares list, or the operating envelope — an RCA whose only artefact is a report has corrected nothing
Thermal events
Notification duties to the AHJ, insurer, utility and offtaker; a re-ignition watch that delays physical access; and a conclusion that usually reaches every identical unit in the fleet

Why the conclusion is contractually loaded

Four verdicts are available and they have different payers. A warranty defect puts the part on the supplier and, depending on the terms, some of the labour, with any revenue consequence reaching only as far as the availability liquidated damages allow. An excluded event — grid outage, curtailment instruction, force majeure, owner-caused — removes the hours from the availability calculation and leaves nobody liable for them.

Operator error moves the cost to the owner or the O&M contractor and can additionally void a warranty claim through the operating-envelope check. A design fault points at the integrator or the EPC and reaches into defects liability and the performance guarantees. Identical downtime, identical hardware, four different outcomes on the money.

That is why the RCA is not a neutral technical exercise once a claim is attached to it, and why the process rules matter more than the analytical method.

Put them in the contract: who performs the investigation and to what deadline, an evidence-preservation obligation triggered by the event, the owner's or independent engineer's right to attend and to receive the raw data rather than a summary, what happens to the availability accounting while the investigation is open, and an escalation route — expert determination or an agreed third party — for when two competent engineers reach two conclusions that pay different people.

The output should also change something. An RCA that lands on an interval rewrites the preventive-maintenance plan; one that lands on a threshold rewrites the predictive setpoints; one that lands on a part nobody had rewrites the critical-spares list; one that lands on a control setting rewrites the operating envelope and possibly the dispatch strategy. An investigation whose only artefact is a report has identified a cause and corrected nothing.

A thermal event is a different investigation

Once a cell has gone into thermal runaway the obligations stop being contractual and start being regulatory. Notification duties run to the authority having jurisdiction, the insurer, and usually the interconnecting utility and the offtaker, on their own timescales rather than the investigation's.

The fire service conducts its own inquiry with its own authority. The site's emergency response plan governs what happens during and immediately after the event, and the hazard mitigation analysis and UL 9540A test data describe what was supposed to happen — the safety detail belongs to those entries, not this one.

Two consequences shape the investigation itself. Access is delayed, because a re-ignition watch keeps people away from the affected enclosure for a period that is a fire-service and safety decision rather than an engineering one, so telemetry carries disproportionate weight in reconstructing an event whose physical evidence may be both damaged and out of reach.

And the conclusion usually reaches past the one enclosure: a runaway traced to a cell defect, a design assumption or a control behaviour implicates every identical unit on the site and often across the vendor's fleet, which is how a single event turns into a fleet-wide corrective campaign. That is also the moment the availability exclusions get read very carefully, because a broad exclusion for response to safety advisories can move the cost of that campaign from the supplier to the owner.

Common pitfalls

Stopping at the failed part is the classic. The fan is replaced, the ticket is closed, and the filter that killed it is still in the plenum. The tell is a corrective history with the same component appearing on the same enclosure at a suspiciously regular interval — the plant is running an unrecognised preventive maintenance program with the failure as its trigger.

Stopping at operator error is the second, and it is usually a hierarchy problem rather than an analytical one. If a procedure permitted the action, the procedure is the cause; if an interface made the wrong setting easy to enter, the interface is. Naming an individual ends the investigation early and reliably prevents the finding that would have stopped the recurrence.

Single-cause bias is the third. Significant events on a plant this heavily instrumented and this heavily protected are usually a latent condition meeting a trigger — a marginal cooling loop meeting the hottest week, a drifted sense channel meeting a hard discharge — because the protections catch the single-cause cases. An RCA that returns exactly one cause on a complex event has probably found the trigger and missed the condition that made it matter.

Common misconception

The RCA is finished once the failed component has been identified and replaced.

In reality: The failed component is the proximate cause, and replacing it restores the plant without touching whatever produced the failure — so the same event returns on the same interval, which is visible in a corrective history where one component keeps reappearing on one enclosure. On a battery asset the report also has to answer a second question the component alone cannot: whether the event was a warranty defect, an excluded event, operator error or a design fault. That determination decides who pays for the part, the labour and the hours, and it is the one the counterparties will actually argue about.

Visuals & further reading
Go deeper

Root cause analysis, in context.

The Grid-Scale BESS course covers root cause analysis — and the rest of the system — from the ground up, the way it actually gets deployed.

Browse the course