Predictive maintenance PdM
Predictive maintenance is work triggered by a measured condition — or by the trend in one — crossing a threshold, rather than by a date. A battery plant is an unusually good candidate for it, because most of the instrumentation is already installed for other reasons.
The BMS measures every cell's voltage at millivolt resolution and every module's temperature continuously, on the order of 5,000 cells in a 5 MWh-class container, because it needs those numbers to protect the string.
Reading them for degradation trends costs data handling, not sensors. The limit is interpretation rather than instrumentation: condition data says a great deal about the slow, thermally driven, population-wide trends and very little about a cell that is about to fail suddenly, and programs sold as covering both are overselling one of them.
Reviewed August 2026 by Sergey Syrvachev
New to BESS? Start free with the 7-email fundamentals course — no cost, no account.
Threshold, trend, and the line against protection
Two things travel under the same label and they are not the same maturity. Condition-based maintenance measures something and raises work when the measurement crosses a limit — coolant conductivity above a value, a bearing temperature above a value, a cell-voltage spread above a value. Prognostics forecasts when the measurement will cross, and schedules the work before it does.
Almost everything running on battery plants today is the first kind with a trend line drawn through it, which is genuinely useful and is not the same claim as predicting a failure date. Being precise about which one a program does is the difference between a maintenance tool and a slide.
The harder line is against protection. The BMS already enforces limits on every cell, and a rack that trips on over-temperature or under-voltage has not given you a predictive signal — it has protected itself after the maintenance decision was already overdue. Predictive thresholds therefore have to sit inside the protection envelope with their own values, their own hysteresis and their own owner, far enough below the trip point that a crew can be mobilised and a part sourced before the plant takes itself off. A PdM alarm set at the protection limit is a fault log with extra steps.
What data actually exists on a battery plant
The battery side is data-rich by construction. Per-cell voltage at millivolt resolution and per-module temperature come from the BMS, along with string current and the per-group resistance estimates the BMS computes anyway to do its own limit arithmetic. Capacity fade is measured properly by reference performance tests on their contractual interval and inferred between them from partial charge and discharge segments.
Cell-voltage divergence within a string is the earliest imbalance signal available, and balancing duty is the derived one worth watching: a system working harder month over month is measuring a divergence rate, and duty concentrated on the same modules points at a thermal gradient or an outlier cell rather than at ordinary drift.
The balance of plant is thinner but not empty. Converters commonly expose heatsink or power-module temperatures — junction temperature is inferred from a thermal model rather than measured — alongside fan runtime hours, internal cabinet temperature, and coolant loop pressure and temperature on liquid-cooled designs.
The thermal plant offers supply-to-return delta-T and approach temperature against a known heat load, which move before cell temperature does: a fouling coil or a weakening pump shows up as a shrinking delta-T while every cell reading is still comfortably in band. Compressor and pump current draw over a start cycle is available from most modern drives and is a genuine early signal, if anyone historises it.
Thermography is the one that needs a person and the one most often done badly. An infrared survey through a closed cubicle door measures the door, so it needs an IR window designed into the panel or a de-energised inspection that defeats the purpose. And the reading is load-dependent — a thermogram taken at 10% of rated current tells you almost nothing about a joint that overheats at full output. On a battery site, current depends on dispatch, so scheduling a meaningful thermal survey is a conversation with whoever controls the plant, not a line on a maintenance calendar.
The instrumentation is already installed for protection: per-cell voltage at millivolt resolution and per-module temperature, on the order of 5,000 cells in a 5 MWh-class container. The limit is interpretation, not sensors. Off-gas detection is a safety function whose output is an alarm and a shutdown, not a work order.
- Trigger
- A measured condition, or its trend, crossing a threshold set deliberately inside the protection envelope — not a date and not a trip
- Native instrumentation
- Per-cell voltage at millivolt resolution and per-module temperature already exist for protection — on the order of 5,000 cells in a 5 MWh-class container
- Sensing floor (LFP)
- On the ~3.2-3.3 V open-circuit plateau, several percent of SOC spread maps to millivolts, at or below ±5 mV sensing — mid-window spread is not evidence of balance
- Trending rule
- Resistance comparisons are valid only at matched SOC, temperature, pulse duration and current direction, against a recorded commissioning baseline
- Best derived signal
- Balancing duty — rising duty concentrated on the same modules measures a divergence rate and points at a thermal gradient or an outlier cell
- Balance-of-plant leading indicators
- HVAC supply-to-return delta-T and approach temperature, PCS heatsink temperature and fan runtime, compressor and pump start-current signature
- Thermography caveats
- An IR survey through a closed door reads the door, and a scan at low load misses a joint that overheats at full current — it needs an IR window and a dispatch conversation
- Not a maintenance function
- Off-gas and thermal-runaway precursor detection is a safety system whose output is an alarm and a shutdown, not a work order
- Fleet limit
- Anomaly models need a population; one site is one climate, one duty and one build batch, so cross-fleet capability sits with the OEM or integrator platform
- Contract dependency
- Owner access to raw BMS and EMS history at a stated resolution and retention — the availability historian at ~1-5 min is fine for trends and blind to transients
Reading the battery trends honestly
Resistance growth is the most valuable trend and the easiest to corrupt. Comparisons are only meaningful at matched state of charge, temperature, pulse duration and current direction; a February rack measurement set against an August commissioning baseline shows dramatic apparent growth that is weather rather than aging.
Either hold the reference conditions for the life of the plant or correct to them, and anchor everything to a real baseline — factory-acceptance DCIR sampling on cells, commissioning impedance records at rack and string level. Trend data assembled from measurements taken wherever the plant happened to be sitting is noise with a slope through it.
Cell-voltage divergence carries a chemistry-specific trap. On LFP's flat open-circuit plateau at about 3.2-3.3 V, cells several percent apart in state of charge read millivolts apart, at or below the ±5 mV accuracy of typical sensing — so a tight mid-window spread on the monitoring screen proves nothing. The honest reads are rest voltages near the ends of the window, the BMS's own charge accounting, and which cells terminate each cycle. Build a PdM threshold on a mid-window spread and you have built an alarm that fires only after the spread is large enough to cost real energy.
Averages are the standing failure of every fleet dashboard. Growth and divergence are never uniform, and the group aging fastest runs hottest under the same current, sags first, and decides when the discharge ends. A widening spread across racks is a finding in its own right while the mean still looks fine, so the trended quantity should be the distribution — worst rack, spread, and the rate at which the tail is moving — not the fleet number.
What it does not predict
A fleet-scale anomaly model needs a fleet. A single site has hundreds of racks and exactly one climate, one duty profile, one commissioning batch and one installation crew, so an unsupervised model trained on it learns that site's habits and calls anything unfamiliar an anomaly. Cross-site models are genuinely more capable, which is why the integrators and OEMs with large installed bases can offer something an owner of two plants cannot build alone. That capability sits behind their platform, which makes access to it a procurement question rather than an engineering one.
Thermal-runaway precursor detection is not a maintenance program and should not be sold inside one. Off-gas, smoke and rapid voltage-collapse detection exist to raise an alarm and initiate a shutdown on a timescale where no work order is relevant; they belong to the safety stack alongside deflagration venting and suppression, and their design, testing and response sit with the emergency response plan and the thermal-runaway hazard chain.
Nothing in the routine telemetry reliably anticipates an internal short from a latent manufacturing defect, which develops far faster than any trend the historian can resolve. That gap is precisely why the safety systems exist in parallel rather than as an output of the monitoring platform.
None of it works without data rights. The historian resolution the availability calculation runs on — commonly around one to five minutes, per contract — is adequate for trending and blind to anything transient, and the raw per-cell BMS history is normally on the supplier's platform rather than the owner's. Write owner access to raw data at a stated resolution and retention into the service agreement, because it is simultaneously the input to every trend on this page and the earliest evidence available in a warranty dispute.
Condition monitoring on a battery plant predicts cell failures before they happen.
In reality: It predicts the slow, thermally driven, population-level trends well — resistance growth, capacity fade, widening divergence, a cooling loop losing capability — because those develop over months and leave a signature in data the BMS already collects. It does not reliably anticipate a sudden internal short from a latent manufacturing defect, which develops on a timescale no historian resolves and no trend line reaches. That is the reason the fire-safety stack exists in parallel with the monitoring platform rather than as an output of it, and why off-gas detection is designed as a safety function rather than a maintenance signal.
- Resistance growth Glossary
- Cell imbalance Glossary
- Why Batteries Fade: The Degradation Mechanisms Behind Every Warranty Table Article
- Interactive: BMS Three-Layer Structure Interactive visual · bess.engineer
Predictive maintenance, in context.
The Grid-Scale BESS course covers predictive maintenance — and the rest of the system — from the ground up, the way it actually gets deployed.