Root cause analysis RCA
Root cause analysis is the structured investigation that asks why a component failed rather than which component failed. The difference is clearest on a real chain. A battery module derates and then trips on high cell temperature. The proximate cause is a failed fan in that container's cooling loop.
The root cause is the filter that sat two quarters past its service interval and loaded the motor, or a control setpoint that let the block run at a combination of duty and ambient the cooling was never sized for. Replace the fan and the plant reproduces the fault on the same schedule.
On this asset class the investigation carries commercial weight as well, because its conclusion sorts the event into warranty defect, excluded event, operator error or design fault — four verdicts with four different payers.
Reviewed August 2026 by Sergey Syrvachev
New to BESS? Start free with the 7-email fundamentals course — no cost, no account.
Proximate cause, root cause, and the layers between
Work the temperature trip down layer by layer and the ownership changes at almost every step. The module tripped because the BMS enforced its limit, which is the protection working. The cells got hot because heat rejection fell. Heat rejection fell because a circulation fan stopped. The fan stopped because its motor was drawing above rating against a loaded filter, or because the bearing reached the end of its life at the hours it was always going to reach it.
The filter was loaded because the interval came from a manual written for a benign site and was applied unchanged to a dusty one, or because the visit was deferred. Each of those sentences has a different owner: the OEM, the O&M contractor, the asset manager who approved the deferral, the designer who picked the interval.
The chain also branches, and taking the first branch on offer is how investigations go wrong. In the same event the fan may have been healthy and the block simply run harder and hotter than the cooling was specified for — which makes it a dispatch and setpoint question rather than a maintenance one, and moves it from the O&M ledger to the design or operating-envelope ledger. Test both branches against the data before committing to one. A cooling loop that has been marginal on hot afternoons all summer leaves a trace in the delta-T record whether or not a fan failed on the day.
Stop when the next why leaves the boundary of what the project can change. An analysis that terminates at the ambient was hot has stopped a layer early: ambient is an input, and the actionable cause is the specification, the setpoint or the interval that assumed a different one. Method matters less than people think — five whys, a fault tree, a cause-and-effect diagram all get you down the same ladder — and the discipline that actually separates a useful RCA from a filed one is whether every step is supported by a record rather than by a recollection.
Evidence has a shelf life
The first action on a significant event is a data freeze, before anything is reset, restarted or cleared. The reason is resolution, and the plant carries several different ones. The historian the availability calculation runs on is commonly around one to five minutes per the contract-defined methodology — fine for downtime accounting and completely blind to a converter event that lasted 200 milliseconds.
The records that resolve that event live in the PCS fault buffer, the protection relay's disturbance recorder and the BMS's high-rate log, all of which are ring buffers of finite depth. They roll. A plant restarted quickly to recover revenue can overwrite the only description of what happened, and no amount of later analysis puts it back.
Pull the whole set, not the obvious one: converter fault buffers and event logs, relay disturbance records, per-cell BMS voltage and temperature history around the event, HVAC and coolant trends for the preceding days, the SCADA event and alarm sequence with its timestamps, and the dispatch record showing what the plant was being asked to do.
Timestamp alignment is the recurring practical failure — devices on different clock sources put the same event in a different order, and a sequence-of-events reconstruction built on drifting clocks produces confident nonsense about which thing caused which.
Physical evidence has its own rules. Where the contract provides for failure analysis, the removed unit goes back intact and tagged rather than into a skip, and the site keeps photographs of the as-found condition before anything is disturbed. There is a genuine tension here between restoring revenue and preserving the record, and it should be resolved in the O&M procedures in advance — a written rule about what gets captured before a restart, with a named person able to authorise the restart, beats the argument that otherwise happens at 3 a.m. with money on the clock.
Freeze the data before any reset: PCS fault buffers, relay disturbance records, per-cell BMS history and SCADA event sequences live in finite ring buffers that roll on restart. An RCA whose only artefact is a report has corrected nothing — it must change an interval, a threshold, the spares list or the operating envelope.
- What it finds
- The underlying cause, not the failed component — the fan is the proximate cause; the overdue filter, or the setpoint that let the block run hot, is the root
- Stopping rule
- Stop at the deepest cause the project can change. A chain that ends at "the ambient was hot" stopped one layer early — ambient is an input, the specification is the cause
- First action
- Freeze the data before any reset or restart: PCS fault buffers, relay disturbance records, per-cell BMS history, HVAC trends, SCADA event sequence and the dispatch record
- Resolution mismatch
- The availability historian at ~1-5 min cannot reconstruct a converter event lasting 200 ms; those records live in finite ring buffers that roll on restart
- Timestamps
- Devices on different clock sources reorder the sequence of events — a reconstruction built on unsynchronised clocks is confidently wrong about causation
- Four verdicts, four payers
- Warranty defect, excluded event, operator error, design fault — identical downtime, entirely different money
- Contract hooks
- Named investigator and deadline, evidence-preservation obligation, owner or IE attendance and raw-data rights, availability treatment while open, and an escalation route for conflicting conclusions
- What it must change
- An interval (PM), a threshold (PdM), the spares list, or the operating envelope — an RCA whose only artefact is a report has corrected nothing
- Thermal events
- Notification duties to the AHJ, insurer, utility and offtaker; a re-ignition watch that delays physical access; and a conclusion that usually reaches every identical unit in the fleet
Why the conclusion is contractually loaded
Four verdicts are available and they have different payers. A warranty defect puts the part on the supplier and, depending on the terms, some of the labour, with any revenue consequence reaching only as far as the availability liquidated damages allow. An excluded event — grid outage, curtailment instruction, force majeure, owner-caused — removes the hours from the availability calculation and leaves nobody liable for them.
Operator error moves the cost to the owner or the O&M contractor and can additionally void a warranty claim through the operating-envelope check. A design fault points at the integrator or the EPC and reaches into defects liability and the performance guarantees. Identical downtime, identical hardware, four different outcomes on the money.
That is why the RCA is not a neutral technical exercise once a claim is attached to it, and why the process rules matter more than the analytical method.
Put them in the contract: who performs the investigation and to what deadline, an evidence-preservation obligation triggered by the event, the owner's or independent engineer's right to attend and to receive the raw data rather than a summary, what happens to the availability accounting while the investigation is open, and an escalation route — expert determination or an agreed third party — for when two competent engineers reach two conclusions that pay different people.
The output should also change something. An RCA that lands on an interval rewrites the preventive-maintenance plan; one that lands on a threshold rewrites the predictive setpoints; one that lands on a part nobody had rewrites the critical-spares list; one that lands on a control setting rewrites the operating envelope and possibly the dispatch strategy. An investigation whose only artefact is a report has identified a cause and corrected nothing.
A thermal event is a different investigation
Once a cell has gone into thermal runaway the obligations stop being contractual and start being regulatory. Notification duties run to the authority having jurisdiction, the insurer, and usually the interconnecting utility and the offtaker, on their own timescales rather than the investigation's.
The fire service conducts its own inquiry with its own authority. The site's emergency response plan governs what happens during and immediately after the event, and the hazard mitigation analysis and UL 9540A test data describe what was supposed to happen — the safety detail belongs to those entries, not this one.
Two consequences shape the investigation itself. Access is delayed, because a re-ignition watch keeps people away from the affected enclosure for a period that is a fire-service and safety decision rather than an engineering one, so telemetry carries disproportionate weight in reconstructing an event whose physical evidence may be both damaged and out of reach.
And the conclusion usually reaches past the one enclosure: a runaway traced to a cell defect, a design assumption or a control behaviour implicates every identical unit on the site and often across the vendor's fleet, which is how a single event turns into a fleet-wide corrective campaign. That is also the moment the availability exclusions get read very carefully, because a broad exclusion for response to safety advisories can move the cost of that campaign from the supplier to the owner.
Common pitfalls
Stopping at the failed part is the classic. The fan is replaced, the ticket is closed, and the filter that killed it is still in the plenum. The tell is a corrective history with the same component appearing on the same enclosure at a suspiciously regular interval — the plant is running an unrecognised preventive maintenance program with the failure as its trigger.
Stopping at operator error is the second, and it is usually a hierarchy problem rather than an analytical one. If a procedure permitted the action, the procedure is the cause; if an interface made the wrong setting easy to enter, the interface is. Naming an individual ends the investigation early and reliably prevents the finding that would have stopped the recurrence.
Single-cause bias is the third. Significant events on a plant this heavily instrumented and this heavily protected are usually a latent condition meeting a trigger — a marginal cooling loop meeting the hottest week, a drifted sense channel meeting a hard discharge — because the protections catch the single-cause cases. An RCA that returns exactly one cause on a complex event has probably found the trigger and missed the condition that made it matter.
The RCA is finished once the failed component has been identified and replaced.
In reality: The failed component is the proximate cause, and replacing it restores the plant without touching whatever produced the failure — so the same event returns on the same interval, which is visible in a corrective history where one component keeps reappearing on one enclosure. On a battery asset the report also has to answer a second question the component alone cannot: whether the event was a warranty defect, an excluded event, operator error or a design fault. That determination decides who pays for the part, the labour and the hours, and it is the one the counterparties will actually argue about.
- Emergency response plan Glossary
- Thermal runaway Glossary
- Availability Glossary
- Interactive: Plant Control Command Path Interactive visual · bess.engineer
Root cause analysis, in context.
The Grid-Scale BESS course covers root cause analysis — and the rest of the system — from the ground up, the way it actually gets deployed.