Commercial

Corrective maintenance CM

Corrective maintenance is the work that restores function after something has already failed or been found defective. Its content is unremarkable: on a battery plant it is almost entirely replacement rather than repair, since modules, racks, PCS power stacks and HVAC compressors are swapped as units and shipped back rather than opened on site. What earns it a page is the clock and the ledger.

Corrective work is the repair half of the availability identity, A ≈ MTBF / (MTBF + MTTR), so every hour of it is deducted from the same annual downtime budget the guarantee is written against. And it is where the commercial machinery engages, because a corrective work order on warranted equipment is also a warranty claim, and who pays for the part, the labour and the lost megawatt-hours is settled by three different documents.

Reviewed August 2026 by Sergey Syrvachev

New to BESS? Start free with the 7-email fundamentals course — no cost, no account.

The clock starts at the fault

Corrective duration is a chain, and only part of it is spannering: detection, diagnosis, mobilisation, parts, the repair itself, then retest and return to service. Mean time to repair conventionally covers the hands-on restoration, and logistics delay — waiting for a spare to arrive — is usually tracked separately, which is exactly why a site can post an excellent repair time and terrible availability at the same time.

A plant that fails rarely and then waits six weeks for a transformer scores worse than one that fails more often and recovers in hours. If your reporting only shows MTTR, it is showing you the part of the outage you already control.

The last link is the one most often left out of the estimate. Return to service on a battery is not a breaker close: the rack has to be reconnected under the BMS, addressed correctly, brought back inside its voltage and temperature limits, and in most cases given time to re-converge with the rest of the string before the block will accept full power.

A replacement that goes in on Tuesday afternoon may not be contributing rated output until later in the week, and whether those hours count as available is a definitional question rather than an engineering one. The threshold rule for partial capability in the availability formula decides it.

The module and rack swap

The workflow is fixed and short. Isolate the rack, open its DC contactors, apply lock-out and tag-out, verify the isolation, break the string connections, remove the failed module, install the replacement, reconnect the sense and communication harness, re-address the module under the BMS, close up, re-balance, return to service.

The step with no analogue elsewhere on the plant is the verification, because a battery module cannot be de-energised. Opening a contactor removes the module from the circuit; it does not remove the energy, and the terminals are live for the rest of the module's life. Insulated tooling, single-hand working and a procedure written for a permanently energised source are the baseline, not an enhancement.

The harness is where the quiet defects live. A mis-mapped or drifted voltage-sense channel gives the BMS a wrong picture of a cell it is supposed to be protecting, and the consequence appears months later as an unexplained outlier or a contested warranty claim rather than as an immediate fault.

Verify the sense mapping after the swap by driving a known condition and confirming the BMS reports it on the expected channel, and record the serial numbers going out and coming in. Module serials are what the OEM's failure analysis, the warranty register and the eventual end-of-life inventory are all keyed to.

A parts warranty covers the part — and only the part.
the PARTcovered by the equipment warrantyLABOUR, freight, crane,mobilisationseparately negotiated — on a remote siteit can cost more than the moduleLOST REVENUEonly via availability liquidateddamages, capped at a share of the annualfee — beyond the cap it returns to theownerone failed modulea warranty claim and a work order at thesame timeAnd the claim still has to pass an envelope check against logged throughput, temperature andmaintenance history before any of it is paid.

Corrective work on a battery is replacement rather than repair: modules, racks, PCS power stacks and HVAC compressors are swapped as units. A fresh module cannot raise an aged string's usable energy, since the old cells still reach their limits first, and it hands the balancing system a permanently wider spread. A module also cannot be de-energised — opening the contactors leaves its terminals live.

Key facts
Trigger
A failure or a defect finding — the work order exists because something is already wrong
What it drives
Corrective duration is the repair half of the availability identity A ≈ MTBF / (MTBF + MTTR) for a repairable system
The hidden half of the outage
Logistics delay — waiting for a spare — is commonly tracked outside MTTR, so excellent repair times can sit behind poor availability
Battery-specific hazard
A module cannot be de-energised. Opening the contactors removes it from the circuit and leaves the terminals live for the rest of its life
Replacement, not repair
Modules, racks, PCS power stacks and HVAC compressors are swapped as units; the failed item goes back for analysis or disposal
New into old
A fresh module cannot raise a string's usable energy — the aged cells still reach the limits first — and it hands the balancing system a permanently wider spread
What balancing can and cannot fix
State-of-charge offset is recoverable at tens to low hundreds of milliamps against hundreds of amp-hours; capacity and resistance mismatch is hardware and is not
PCS configuration risk
A replacement stack carries factory parameters — the site's grid-code parameter and firmware file must be reloaded as a defined return-to-service step
Three costs, three mechanisms
Part under the equipment warranty; labour, freight and access separately negotiated; lost revenue only via capped availability liquidated damages, if at all
Same event, two ledgers
An outage can be a supplier defect and an excused availability period at once — and which of the four verdicts applies is the RCA's output

A new module in an aged string

Dropping a factory-fresh module into a string that has been cycling for six years creates a problem the BMS then has to carry for the rest of the asset's life. The new cells have more capacity and lower internal resistance than their neighbours.

Every cell in the series string carries the same current, so the aged cells still change state of charge faster per amp-hour moved, still reach the voltage limits first, and still decide when the charge and the discharge end. The new module's extra capacity is unreachable. What the swap has actually bought is a working module in place of a broken one, not a larger string.

What it has also bought is a permanently wider spread for the balancing system to manage. Passive bleed balancing moves tens to low hundreds of milliamps against cells of hundreds of amp-hours, so re-convergence is measured in hours to days and continues indefinitely, dissipating energy as heat inside the enclosure.

And only part of the spread is addressable at all: the state-of-charge offset can be balanced out, while the difference in capacity and resistance between new and aged cells is hardware and no balancing current touches it. The mismatch also runs slightly against the usual thermal gradient, since the low-resistance new module dissipates less I²R heat than the cells around it under identical current.

The practical answers are all forms of matching. OEMs commonly hold capacity-matched or age-matched replacement stock for exactly this reason, and where a string is well into life a rack-level replacement can be the cleaner intervention than a module-level one. Adding capacity, as opposed to restoring it, is a different problem with a different rule: vendors typically require augmentation racks on a separate DC bus or a dedicated PCS input rather than paralleled onto an aged bus, because paralleling fresh and faded strings forces the pack to the weakest string's window.

PCS power-stage replacement

On a modular converter the replaceable unit is the power stack or power module, and the granularity of the design decides how much of the plant the fault takes with it — one stack out of several in a block is a partial derate, while a fault that strands the whole block is a full one. That difference is exactly the difference between a time-based availability metric reading near 100% and a capacity-weighted one reporting the true pro-rata loss, so the same corrective event scores very differently depending on which metric the contract uses.

The configuration travels with the hardware, and this is the step that catches out crews who have done the mechanical work correctly. A replacement stack arrives with the factory's parameter set, and the site is running a grid-code parameter file that was tuned, tested and in some jurisdictions witnessed: ride-through envelopes, droop settings, reactive priority, protection thresholds.

Restoring the block without restoring its parameter file returns a converter to service that no longer matches what the interconnection was approved against. Keep the as-built parameter and firmware register under change control, load it as a defined step in the return-to-service procedure, and check whether the interconnection agreement requires any re-verification after a change to converter control settings.

The warranty claim behind the work order

Corrective work on warranted equipment splits into three costs that are covered by three different mechanisms. The part is normally covered by the equipment warranty. Labour, freight, crane, site access and the specialist technician's mobilisation are separately negotiated, are frequently not covered by a bare parts warranty, and on a container-scale swap in a remote location can exceed the cost of the part.

Lost revenue is covered by neither: to the extent it is addressed at all, it comes through the availability guarantee's liquidated damages in the long-term service agreement, which are typically capped as a percentage of the annual service fee. A project that assumed the warranty made it whole discovers the gap on its first significant claim.

The claim then has to survive an envelope check. Suppliers assess against the logged operating history — throughput, cell temperature, state-of-charge excursions, maintenance completion — so the BMS record and the preventive-maintenance file are the claim, and the technical merits are argued afterwards. Preserve the failed unit where the contract calls for failure analysis, because a claim resting only on telemetry is weaker than one with the hardware attached to it.

Then there is the interaction that surprises people, which is that a warranty defect and an availability exclusion are not opposites. The same outage can be a supplier defect that triggers a parts and labour obligation and, separately, an excused period under the availability guarantee — many agreements excuse downtime taken for warranty repairs, so the supplier pays for the module and the hours leave the denominator anyway.

Read the two documents side by side, and note which of the four possible verdicts on the event is being asserted: warranty defect, excluded event, operator error or design fault. That determination is the output of the root cause analysis, and it decides which of the three cost buckets anyone is arguing about.

Common misconception

The equipment is under warranty, so a module failure costs the project nothing.

In reality: A parts warranty covers the part. Labour, freight, crane and specialist mobilisation are separately negotiated and on a remote site can cost more than the module. Lost revenue is not covered by either and is only addressed, if at all, through the availability guarantee's liquidated damages, which are typically capped at a share of the annual service fee — so beyond the cap the shortfall returns to the owner. And the claim still has to pass an envelope check against logged throughput, temperature and maintenance history before any of it is paid.

Visuals & further reading
Go deeper

Corrective maintenance, in context.

The Grid-Scale BESS course covers corrective maintenance — and the rest of the system — from the ground up, the way it actually gets deployed.

Browse the course