Reactive Maintenance Is a Data Problem, Not a Resource Problem 

Ask an engineering manager why their team spends so much time firefighting and the answer is usually resource. Not enough people, not enough hours, too much site to cover. 

That is real. But it isn’t usually the cause. 

Most reactive work happens because the first indication of a problem was the problem itself. A breaker operated. A UPS alarmed. A line stopped. At that point the team responds well, often very well, but they are responding, and the timing was chosen by the fault. 

Adding headcount lets you respond faster. It doesn’t move the moment you find out. 

The interval is the exposure

Scheduled maintenance is not the weak link. The gap around it is.

Manufacturers of LV and MV distribution equipment and transformers typically recommend maintenance every three years under normal environmental conditions, tightening to two in severe ones. That’s a reasonable interval for planned intervention.

It is a long time for an asset to go unobserved.

Between those visits, load grows, new equipment is commissioned, production patterns shift, insulation degrades and supply quality varies. None of it announces itself. IDC research found that 38% of companies reported events resulting in production losses. The events that cause them rarely begin on the day they surface.

So a site can hold a complete, well-executed maintenance record and still have no idea what happened in the 34 months between visits.

What that looks like on site

We were asked to review the HV infrastructure at a UK food manufacturer with a full maintenance history and a competent engineering team.

The circuit breakers had not been maintained since 2008, twelve years overdue. A fault would not have been reliably protected, and the DNO could have forced a shutdown on the grounds of inadequate protection. On the transformer, every core clamp had vibrated loose, and one had made contact with the LV windings.

Nobody had been negligent. The records showed what had been done, not what had changed. A visual inspection during normal operation would not have found a loose core clamp, and there was no continuous measurement that would have flagged the drift.

That is the shape of the problem: missing observation, not missing effort.

Why symptoms send teams to the wrong place

When the only information is the symptom, investigation starts at the symptom. It is usually the wrong starting point.

A UPS alarms repeatedly, and the fault sits upstream in the incoming supply. A panel runs hot, and the cause is harmonic current rather than the panel. At 40% current distortion, heat losses run around 10% higher, equivalent to carrying about 5% more current than the design assumed. A drive trips, and it’s responding to a voltage dip lasting a fraction of a second that nothing on site recorded.

Each of those investigations can be conducted competently and still reach the wrong conclusion, because the evidence that would identify the cause was never captured. The component gets replaced, the symptom clears, and the fault returns, at which point the same investigation happens again.

Repeated investigation of the same fault is the largest hidden cost in reactive maintenance, and it never appears in a maintenance budget as such.

Repeat faults are patterns not incidents

The single most useful question about a recurring fault is not what failed but what else was happening at the time. 

Was the site at peak demand? Did it coincide with a particular machine starting? Were several assets affected in the same second? Was there a voltage or current disturbance on the incoming supply? 

Answering that requires two things a symptom-based approach can’t provide: a time-stamped record across multiple points in the system, and enough history to distinguish a coincidence from a pattern. With them, a scatter of unrelated call-outs often resolves into two or three root causes. 

Without them, you have four separate jobs that will each be done twice. 

The economics are better documented than most teams realise

The case for shifting effort earlier isn’t only about workload. It shows up in cost.

Frost & Sullivan’s analysis of asset performance management found that savings from preventive maintenance, compared with a corrective break-fix approach, can reach 12–18%. McKinsey put a figure on the next step, finding that improved condition monitoring typically reduces maintenance costs by 10–15%.

Both point in a direction that surprises people: condition-based monitoring does not always mean doing more maintenance sooner. Sometimes the evidence shows an asset is stable and intervention can be safely deferred and scheduled around production, rather than performed on a calendar because the calendar said so.

The aim is not more maintenance, but maintenance aimed at the right asset at a time you chose.

What to monitor first

You don’t need coverage of everything, and attempting it is usually how these programmes stall.

Start with the assets whose failure would stop the site or create a safety or compliance problem: typically the incoming supply, main HV and LV switchgear, transformers, UPS systems, cable terminations and the boards feeding critical process. For those, the things worth tracking continuously are asset condition, temperature, load behaviour, power quality events and partial discharge activity.

Then make sure the output reaches the people planning the work. Monitoring that produces a dashboard nobody reviews before the maintenance meeting,has changed nothing. The measure of whether it’s working is simple: does next month’s schedule look different because of what the data showed?

If it doesn’t, you’ve bought instrumentation, not insight.

From insight to resolution

Identifying a developing problem is the start. What usually determines whether reactive work actually falls is what happens next: whether the remedial work, the maintenance and the corrective engineering get delivered, and whether the follow-up measurement confirms the issue was resolved rather than moved.

Acteniq works across that whole path. We monitor critical electrical assets, interpret what the evidence shows, and where internal teams don’t have the time, resource or specialist capability, we deliver the maintenance, remediation and monitoring implementation required, then verify the result against the baseline.

See how one food manufacturer moved from emergency call-outs to planned intervention  

Work with engineers, not an account team.

Sources cited in this piece

  • Schneider Electric, Benefits of shifting from traditional to condition-based maintenance in electrical distribution equipment, 998-22447106_GMA, 2022
  • IDC, Maximize Business and Operation Resiliency through Services
  • Frost & Sullivan, Data-driven Asset Performance Management, 2020
  • McKinsey, Coronavirus: Industrial IoT in challenging times, April 2020
  • Schneider Electric, Masterpact NT and NW Maintenance Guide, LVPED508016EN-02, 07/2013 (harmonic heat loss)
  • Acteniq site findings, UK food manufacturer, onboarding review

Share this

Reactive maintenance often starts with missing visibility, not a lack of resource. See how condition monitoring helps engineers identify recurring patterns, reduce repeat faults and plan maintenance around evidence.

Related Posts