Preventive maintenance PM
Preventive maintenance is work scheduled by elapsed time or by an accumulated counter — months since the last visit, fan hours, compressor hours, breaker operations — and carried out whether or not the equipment shows any sign of needing it. That trigger is the whole definition. Predictive maintenance waits for a measurement to cross a threshold; corrective maintenance waits for a failure; preventive maintenance does not wait.
On a grid-scale battery the recurring list is shorter than on a thermal plant and badly lopsided: HVAC service dominates it, because the thermal system is what holds the cells near their design temperature while nearly everything else on the site is solid-state or bolted together.
The rest is torque verification on the power path, cleaning, code-driven fire-system inspection, firmware campaigns, insulation checks, and the reference capacity test where the contract calls for one.
Reviewed August 2026 by Sergey Syrvachev
New to BESS? Start free with the 7-email fundamentals course — no cost, no account.
The trigger is a date, not a symptom
The three maintenance classes are separated by a single question: what caused the work order to exist. Preventive work is raised by a schedule, and a runtime counter is still a schedule. A fan-hours meter records how long the fan has been asked to run, not how its bearing is doing; the counter is a proxy for wear chosen in advance, and it fires on the unit that needed the visit and the unit that did not, alike.
That is the trade PM makes. It accepts doing some work early in exchange for doing very little work late, and it is the right trade wherever the failure mode is wear-driven, the interval is knowable, and the consequence of finding out the hard way is expensive.
Two sources set the intervals, and they answer to different people. Most come from the manufacturer's operations and maintenance manual, and those are load-bearing well beyond reliability: skipped required maintenance is a listed condition under which a capacity claim can be voided or pro-rated, so the completed PM record is warranty evidence before it is anything else. The other source is code.
Fire detection and suppression equipment carries inspection, testing and maintenance regimes written into the standards that govern each device — NFPA 72 for detection and alarm, NFPA 25 where the suppression is water-based, NFPA 2001 for clean-agent systems — rather than into the battery vendor's manual. Intervals that come from a code do not bend around a dispatch schedule, and the authority having jurisdiction accepted the installation on the assumption that they would be kept.
Some counters are written by the trading desk rather than by the calendar. An on-load tap changer's maintenance is counted in operations, not years, and on a plant that reverses between charge and discharge the operation count follows dispatch — a plant cycling once a day passes through the reversal twice, and one chasing a regulation signal moves far more often.
Medium-voltage circuit breakers are the same shape, classified by mechanical and electrical endurance in operations. So part of the substation PM plan is set by how the asset is being traded, which means the maintenance budget moves when the revenue strategy does.
HVAC service, and what it costs when it slips
Filters, condenser and evaporator coils, coolant level and concentration on liquid loops, pump and fan checks, condensate drains and the refrigerant circuit make up the single largest block of recurring work on a battery site. It is also the block whose neglect is hardest to see, because it does not present as a failure.
A loaded filter or a fouled coil reduces heat rejection; cell temperature climbs toward the top of the roughly 15-35 °C LFP window instead of sitting near the 25 °C design target; and the battery management system derates current before anything trips. The plant is up, the outage log is clean, and the fleet is aging faster than the model assumed.
The price is paid on the degradation curve. Fade rate roughly doubles per 10 °C of sustained cell temperature, so a summer spent a few degrees hot is bought out of the warranted retention rather than out of the maintenance budget, and it surfaces at the next reference performance test as a shortfall nobody can attribute.
Intra-container gradient does the same damage unevenly: liquid-cooled enclosures commonly specify cell-to-cell spread of 3 °C or tighter and air-cooled designs around 5 °C, and a partially blocked airflow path widens that spread in one physical location, writing a permanent imbalance into the cells that sit there.
Filter intervals are the most site-specific numbers in the whole plan and the most often copied from somewhere else. Dust, pollen, salt air, agricultural season and the height of the intake above grade all move the real interval, and none of them is in the OEM manual. Set the first year's interval from the manual, then measure — differential pressure across the filter bank, or supply-to-return temperature difference against a known heat load — and let the site tell you what its interval actually is.
HVAC service dominates the recurring list. Fire detection and suppression carry inspection intervals set by code — NFPA 72, 25 and 2001 each with their own regime — which apply whether or not anything has moved. Battery-system PM sits in the LTSA and balance-of-plant PM in the O&M agreement: two schedules needing one outage window per block.
- Trigger
- Elapsed time or an accumulated counter (fan hours, compressor hours, breaker or tap-changer operations) — never a measured condition of the specific unit
- Dominant recurring task
- HVAC service — filters, coils, coolant, pumps and drains; neglect presents as thermal derate and accelerated fade, not as an outage
- Why thermal slip is expensive
- Fade rate roughly doubles per 10 °C of sustained cell temperature, so a hot season is paid out of the warranted retention curve rather than the maintenance budget
- Two sources of intervals
- The OEM manual (warranty-conditioning) and the codes governing each fire-protection device — NFPA 72, NFPA 25, NFPA 2001 each carry their own inspection, testing and maintenance regime
- Insulation on the DC side
- A 1500 VDC array is normally ungrounded and monitored continuously by an insulation monitoring device; the PM task is proving the monitor works, with offline IR testing only on isolated circuits
- Contractual status of a PM window
- Negotiated — excluded outright, capped in hours, counted in full, or excluded only when scheduled with notice. There is no standard treatment
- Downtime budget for context
- 98% availability ≈ 175 h/yr of unexcused downtime, 97% ≈ 263 h/yr, 95% ≈ 438 h/yr
- Warranty link
- Skipped required maintenance is a listed condition under which a capacity claim can be voided or pro-rated — the completed PM record is the evidence
- Scope split
- Battery-system PM in the LTSA with the OEM, balance-of-plant PM in the O&M agreement — two schedules that need coordinating onto one outage window per block
- Also on the calendar
- The reference performance test where the contract requires one (commonly annual or biennial) — it consumes warranted throughput and takes the block out of the market
The rest of the list
Bolted joints in the power path relax under thermal cycling: module terminals, rack busbars, DC combiner lugs, PCS terminations and medium-voltage cable lugs. A relaxed joint carries more contact resistance, dissipates more I²R heat at the interface, and relaxes further, so the failure mode compounds on itself.
The PM task is verification against the manufacturer's stated torque with a calibrated wrench, with the reading recorded rather than a tick placed in a box. Read the procedure carefully on one point: whether it asks you to check a joint or to break and re-tighten it. Repeatedly re-torquing plated and aluminium joints has real detractors, and the OEM's written instruction is what the warranty is measured against either way.
Firmware campaigns arrive on the vendor's release schedule rather than yours, and they take the block offline. BMS controllers, PCS control cards, the plant controller and the site's protection relays each carry their own release trains, and a campaign is only finished when the as-built firmware register matches what is actually running on every unit — a register that drifts quietly every time a board is swapped under a corrective work order.
Insulation is the other switched-off item. A 1500 VDC array is normally ungrounded and watched continuously by an insulation monitoring device, so the routine task is proving the monitor still works rather than taking the measurement; offline insulation-resistance testing is done on isolated cables and converter circuits with the strings disconnected, because you cannot apply a test voltage across a live battery.
Where the contract requires a reference performance test — commonly annually or every two years — it lands on the PM calendar even though it is a measurement rather than a service. It takes the block out of the market for the duration, it consumes throughput that counts against the warranty's annual allowance, and it produces the compliance record the capacity guarantee is argued from.
Protection relays get the same treatment: secondary injection and functional trip testing on an interval, because a relay that has never been tested since commissioning is an assumption rather than a protection scheme.
Whether a PM window counts against availability
There is no default, and this is the clause to find before anyone talks about a maintenance plan. Some agreements exclude scheduled maintenance outright. Some cap it at an agreed number of hours a year and count the overrun. Some count it in full and expect the supplier to work around dispatch.
Some exclude it only where the outage was noticed and scheduled inside an agreed window, which turns an unscheduled overrun into a countable one. Do the arithmetic before conceding the point: at 98% the entire annual unexcused-downtime allowance is about 175 hours, at 97% about 263, and at 95% about 438. Where planned work is counted in full against that budget, the maintenance plan and the failure rate are drawing on the same account.
The scope split matters as much as the exclusion. Battery-system PM normally sits with the OEM under the long-term service agreement while the balance of plant sits with the O&M contractor, so two organisations with two schedules and two definitions of a maintenance outage arrive at the same substation. Coordinate them onto the same window per block, or the plant pays for the same downtime twice and the availability report has two authors with different arithmetic.
Firmware is the awkward case. Watch for an exclusion covering response to safety advisories, which reads as routine housekeeping and can quietly excuse a fleet-wide corrective campaign — a supplier defect, dressed as scheduled maintenance, removed from the denominator. The clean version enumerates what a scheduled maintenance exclusion covers, caps the hours, and puts anything raised by a defect finding on the corrective side of the ledger where it belongs.
Common pitfalls
Year one is where PM programs are lost. The site is still working through a commissioning punch list, crews are already on the ground doing other things, and the first scheduled service slips because everything looks new. Infant-mortality faults are also concentrated in that period, so the year the plant most needs a clean baseline of torque readings, filter conditions and insulation values is the year it is least likely to get one.
The second failure is the record. A visit performed and not documented is, at claim time, a visit that did not happen — the envelope check reads logged evidence, and a supplier assessing a capacity claim will ask for the maintenance history before it asks anything else. Photographs, measured values and the technician's name beat a completion checkbox, and the file has to survive an O&M contractor changing hands, which it frequently does not.
The third is servicing only what alarms. A PM program that has quietly reorganised itself around whichever containers raised a ticket last month has stopped being preventive and become corrective work with a schedule attached. The containers that never alarm are also the ones nobody has opened in two years.
A battery plant has no moving parts, so it barely needs preventive maintenance.
In reality: The cells have no moving parts. Everything keeping them alive does — fans, pumps, compressors, dampers, filters and contactors — and the thermal system is both the busiest mechanical plant on site and the one whose degradation is invisible in the outage log. A container that quietly loses heat-rejection capacity never trips; it derates, runs a few degrees warm, and pays for it on the degradation curve. Separately, the fire detection and suppression equipment carries inspection intervals set by code rather than by the vendor, and those apply whether or not anything has moved.
- HVAC Glossary
- Availability Glossary
- BESS Commissioning: How a Container Full of Cells Becomes a Power Plant Article
- Interactive: BESS Container Structure Interactive visual · bess.engineer
Preventive maintenance, in context.
The Grid-Scale BESS course covers preventive maintenance — and the rest of the system — from the ground up, the way it actually gets deployed.