Commercial

Predictive maintenance PdM

Predictive maintenance is work triggered by a measured condition — or by the trend in one — crossing a threshold, rather than by a date. A battery plant is an unusually good candidate for it, because most of the instrumentation is already installed for other reasons.

The BMS measures every cell's voltage at millivolt resolution and every module's temperature continuously, on the order of 5,000 cells in a 5 MWh-class container, because it needs those numbers to protect the string.

Reading them for degradation trends costs data handling, not sensors. The limit is interpretation rather than instrumentation: condition data says a great deal about the slow, thermally driven, population-wide trends and very little about a cell that is about to fail suddenly, and programs sold as covering both are overselling one of them.

Reviewed August 2026 by Sergey Syrvachev

New to BESS? Start free with the 7-email fundamentals course — no cost, no account.

Threshold, trend, and the line against protection

Two things travel under the same label and they are not the same maturity. Condition-based maintenance measures something and raises work when the measurement crosses a limit — coolant conductivity above a value, a bearing temperature above a value, a cell-voltage spread above a value. Prognostics forecasts when the measurement will cross, and schedules the work before it does.

Almost everything running on battery plants today is the first kind with a trend line drawn through it, which is genuinely useful and is not the same claim as predicting a failure date. Being precise about which one a program does is the difference between a maintenance tool and a slide.

The harder line is against protection. The BMS already enforces limits on every cell, and a rack that trips on over-temperature or under-voltage has not given you a predictive signal — it has protected itself after the maintenance decision was already overdue. Predictive thresholds therefore have to sit inside the protection envelope with their own values, their own hysteresis and their own owner, far enough below the trip point that a crew can be mobilised and a part sourced before the plant takes itself off. A PdM alarm set at the protection limit is a fault log with extra steps.

What data actually exists on a battery plant

The battery side is data-rich by construction. Per-cell voltage at millivolt resolution and per-module temperature come from the BMS, along with string current and the per-group resistance estimates the BMS computes anyway to do its own limit arithmetic. Capacity fade is measured properly by reference performance tests on their contractual interval and inferred between them from partial charge and discharge segments.

Cell-voltage divergence within a string is the earliest imbalance signal available, and balancing duty is the derived one worth watching: a system working harder month over month is measuring a divergence rate, and duty concentrated on the same modules points at a thermal gradient or an outlier cell rather than at ordinary drift.

The balance of plant is thinner but not empty. Converters commonly expose heatsink or power-module temperatures — junction temperature is inferred from a thermal model rather than measured — alongside fan runtime hours, internal cabinet temperature, and coolant loop pressure and temperature on liquid-cooled designs.

The thermal plant offers supply-to-return delta-T and approach temperature against a known heat load, which move before cell temperature does: a fouling coil or a weakening pump shows up as a shrinking delta-T while every cell reading is still comfortably in band. Compressor and pump current draw over a start cycle is available from most modern drives and is a genuine early signal, if anyone historises it.

Thermography is the one that needs a person and the one most often done badly. An infrared survey through a closed cubicle door measures the door, so it needs an IR window designed into the panel or a de-energised inspection that defeats the purpose. And the reading is load-dependent — a thermogram taken at 10% of rated current tells you almost nothing about a joint that overheats at full output. On a battery site, current depends on dispatch, so scheduling a meaningful thermal survey is a conversation with whoever controls the plant, not a line on a maintenance calendar.

Condition monitoring reads months, not milliseconds.
what trending SEESresistance growth · capacity fade · wideningdivergence, with balancing duty concentrating on thesame modules · a cooling loop losing capability — alldeveloping over months, in data the BMS alreadycollectswhat it CANNOT seea sudden internal short from a latent manufacturingdefect — a timescale no historian resolves and notrend line reachesWhich is why the fire-safety stack exists in PARALLEL with the monitoring platform, rather thanas an output of it.

The instrumentation is already installed for protection: per-cell voltage at millivolt resolution and per-module temperature, on the order of 5,000 cells in a 5 MWh-class container. The limit is interpretation, not sensors. Off-gas detection is a safety function whose output is an alarm and a shutdown, not a work order.

Key facts
Trigger
A measured condition, or its trend, crossing a threshold set deliberately inside the protection envelope — not a date and not a trip
Native instrumentation
Per-cell voltage at millivolt resolution and per-module temperature already exist for protection — on the order of 5,000 cells in a 5 MWh-class container
Sensing floor (LFP)
On the ~3.2-3.3 V open-circuit plateau, several percent of SOC spread maps to millivolts, at or below ±5 mV sensing — mid-window spread is not evidence of balance
Trending rule
Resistance comparisons are valid only at matched SOC, temperature, pulse duration and current direction, against a recorded commissioning baseline
Best derived signal
Balancing duty — rising duty concentrated on the same modules measures a divergence rate and points at a thermal gradient or an outlier cell
Balance-of-plant leading indicators
HVAC supply-to-return delta-T and approach temperature, PCS heatsink temperature and fan runtime, compressor and pump start-current signature
Thermography caveats
An IR survey through a closed door reads the door, and a scan at low load misses a joint that overheats at full current — it needs an IR window and a dispatch conversation
Not a maintenance function
Off-gas and thermal-runaway precursor detection is a safety system whose output is an alarm and a shutdown, not a work order
Fleet limit
Anomaly models need a population; one site is one climate, one duty and one build batch, so cross-fleet capability sits with the OEM or integrator platform
Contract dependency
Owner access to raw BMS and EMS history at a stated resolution and retention — the availability historian at ~1-5 min is fine for trends and blind to transients

Resistance growth is the most valuable trend and the easiest to corrupt. Comparisons are only meaningful at matched state of charge, temperature, pulse duration and current direction; a February rack measurement set against an August commissioning baseline shows dramatic apparent growth that is weather rather than aging.

Either hold the reference conditions for the life of the plant or correct to them, and anchor everything to a real baseline — factory-acceptance DCIR sampling on cells, commissioning impedance records at rack and string level. Trend data assembled from measurements taken wherever the plant happened to be sitting is noise with a slope through it.

Cell-voltage divergence carries a chemistry-specific trap. On LFP's flat open-circuit plateau at about 3.2-3.3 V, cells several percent apart in state of charge read millivolts apart, at or below the ±5 mV accuracy of typical sensing — so a tight mid-window spread on the monitoring screen proves nothing. The honest reads are rest voltages near the ends of the window, the BMS's own charge accounting, and which cells terminate each cycle. Build a PdM threshold on a mid-window spread and you have built an alarm that fires only after the spread is large enough to cost real energy.

Averages are the standing failure of every fleet dashboard. Growth and divergence are never uniform, and the group aging fastest runs hottest under the same current, sags first, and decides when the discharge ends. A widening spread across racks is a finding in its own right while the mean still looks fine, so the trended quantity should be the distribution — worst rack, spread, and the rate at which the tail is moving — not the fleet number.

What it does not predict

A fleet-scale anomaly model needs a fleet. A single site has hundreds of racks and exactly one climate, one duty profile, one commissioning batch and one installation crew, so an unsupervised model trained on it learns that site's habits and calls anything unfamiliar an anomaly. Cross-site models are genuinely more capable, which is why the integrators and OEMs with large installed bases can offer something an owner of two plants cannot build alone. That capability sits behind their platform, which makes access to it a procurement question rather than an engineering one.

Thermal-runaway precursor detection is not a maintenance program and should not be sold inside one. Off-gas, smoke and rapid voltage-collapse detection exist to raise an alarm and initiate a shutdown on a timescale where no work order is relevant; they belong to the safety stack alongside deflagration venting and suppression, and their design, testing and response sit with the emergency response plan and the thermal-runaway hazard chain.

Nothing in the routine telemetry reliably anticipates an internal short from a latent manufacturing defect, which develops far faster than any trend the historian can resolve. That gap is precisely why the safety systems exist in parallel rather than as an output of the monitoring platform.

None of it works without data rights. The historian resolution the availability calculation runs on — commonly around one to five minutes, per contract — is adequate for trending and blind to anything transient, and the raw per-cell BMS history is normally on the supplier's platform rather than the owner's. Write owner access to raw data at a stated resolution and retention into the service agreement, because it is simultaneously the input to every trend on this page and the earliest evidence available in a warranty dispute.

Common misconception

Condition monitoring on a battery plant predicts cell failures before they happen.

In reality: It predicts the slow, thermally driven, population-level trends well — resistance growth, capacity fade, widening divergence, a cooling loop losing capability — because those develop over months and leave a signature in data the BMS already collects. It does not reliably anticipate a sudden internal short from a latent manufacturing defect, which develops on a timescale no historian resolves and no trend line reaches. That is the reason the fire-safety stack exists in parallel with the monitoring platform rather than as an output of it, and why off-gas detection is designed as a safety function rather than a maintenance signal.

Go deeper

Predictive maintenance, in context.

The Grid-Scale BESS course covers predictive maintenance — and the rest of the system — from the ground up, the way it actually gets deployed.

Browse the course