Meet TAI - Industrial AI grounded in your digital twin
Blog · Maintenance

From reactive to predictive: IIoT maintenance that prevents downtime

Unplanned downtime is not a maintenance problem, it is a business problem with a maintenance cause. Siemens’ analysis of the Global 500 puts the total cost of unplanned stoppages at roughly $1.4 trillion a year, a figure that has risen 62% since 2019 as plants have grown more interconnected and production margins thinner. Behind that number are cascading effects: scrapped materials, delayed shipments, emergency labour premiums, regulatory penalties, and reputational damage with customers who had come to expect continuous supply.

The industries spending the most to prevent those failures, energy, process manufacturing, automotive assembly, heavy mining, are also the industries where maintenance models have changed least. Most still operate somewhere between “fix it when it breaks” and “service it every ninety days whether or not it needs it.” Industrial IoT sensor costs have dropped dramatically, compute has moved to the edge, and machine-learning libraries are commoditised, yet PdM programmes routinely stall in pilot. Understanding why requires a clear picture of what each maintenance strategy actually offers, and what it demands.

The maintenance strategy spectrum

Maintenance philosophy exists on a continuum from reactive at one end to prescriptive at the other. Each stage represents a different relationship between data, decision, and action.

Reactive maintenance

Run-to-failure is the oldest and simplest model: do nothing until the asset stops working, then repair it. Planned reactive maintenance is rational for non-critical, cheap-to-replace components where the cost of an unplanned failure is genuinely lower than the cost of monitoring or scheduled servicing. For most industrial equipment, however, the failure cost is unpredictable and the knock-on consequences make it unacceptable as a primary strategy.

Preventive maintenance

Time-based or usage-based servicing removes the randomness of reactive failure by replacing or inspecting at fixed intervals. It is the dominant strategy across manufacturing and utilities and it works reasonably well when failure modes are strongly time-dependent. The problem is that most equipment failure is not time-dependent in a simple way. Oil degradation, bearing wear, and seal fatigue all vary with operating conditions, load cycles, ambient temperature, and process chemistry. A fixed calendar therefore produces two systematic errors: servicing assets that do not yet need it (wasting parts, labour, and planned downtime) and missing assets that have degraded faster than the interval assumed. According to IBM’s overview of predictive maintenance, up to 30% of scheduled maintenance activities deliver no benefit because the equipment was still within tolerance at service time.

Condition-based maintenance

Condition-based maintenance (CBM) introduces the idea of acting on measured asset state rather than on time. A technician takes a vibration reading or draws an oil sample on a defined schedule, compares the result to acceptance limits, and decides whether to intervene. It is more targeted than calendar-based work, but it still depends on periodic human data collection, and the decision window between readings can be too narrow to prevent failure on fast-degrading components.

Predictive maintenance

Predictive maintenance (PdM) extends CBM by making the condition monitoring continuous and automating the pattern detection. Sensors stream data in near-real time to analytics software that flags early signs of developing failure, often hours, days, or weeks before the fault would become visible to a technician. The intervention can then be planned for a convenient window, preventing both the unplanned event and the unnecessary scheduled service.

Prescriptive maintenance

Prescriptive approaches go one step further: they not only predict when something will fail but recommend the optimal response, weighing part lead times, crew availability, production schedules, and risk levels to generate a prioritised action. Prescriptive tools are maturing in capital-intensive sectors but remain dependent on the quality of the PdM programme beneath them.

What data predictive maintenance actually uses

PdM is not a single technique. Different asset classes fail in different ways, and the sensor signals that reveal those failure modes differ accordingly.

  • Vibration analysis. The workhorse of rotating equipment monitoring. Accelerometers on bearing housings, motor frames, and pump casings capture frequency spectra that reveal bearing defects, imbalance, misalignment, looseness, and gear mesh anomalies. Baseline spectra are captured during healthy operation; deviations from that baseline trigger alerts. The physics is well understood, which is why vibration analysis has the most mature body of thresholds and failure mode libraries.
  • Thermal imaging and point temperature sensors. Infrared cameras and thermocouples detect heat signatures that indicate electrical overload, loose connections, blocked cooling passages, and insulation breakdown. Thermal data is particularly valuable for electrical switchgear, transformers, and refractory linings, where failures progress faster than vibration signatures suggest.
  • Acoustic emission and ultrasound. High-frequency sound sensors detect compressed-air leaks, partial discharge in high-voltage equipment, early-stage bearing defects, and valve seat erosion. Ultrasound is especially useful in noisy environments where vibration analysis is confounded by background frequencies from adjacent equipment.
  • Oil and fluid analysis. Particle counting, viscosity measurement, and spectrometric analysis of lubricant samples reveal wear metals, contamination, oxidation, and coolant ingress. Oil analysis is slower than electronic sensors because samples must be drawn and processed, but it provides chemical information that no surface-mounted sensor can replicate.
  • Motor current signature analysis (MCSA). Current and voltage waveforms at the motor controller encode information about rotor bar condition, eccentricity, and load fluctuations. MCSA requires no additional sensors beyond the drive cabinet instrumentation that already exists in most plants, making it an economical route to monitoring conveyor drives, pumps, and fans.

In practice, mature PdM programmes use more than one modality for critical assets. Vibration alone may catch a bearing defect three weeks before failure; pairing it with oil metal content and thermal data reduces false positives and extends the prediction window.

Models versus thresholds

The technical debate in PdM programme design centres on whether to use simple threshold alerting or machine-learning models. The distinction matters more than most vendors admit.

Threshold-based rules (alert when vibration RMS exceeds 10 mm/s, or when temperature rises more than 15°C above ambient) are transparent, maintainable, and fast to implement. They work well when the failure mode is well characterised and the operating regime is stable. They break down when assets run at variable loads, when ambient conditions change seasonally, or when the failure pattern is subtle and multivariate rather than a simple level crossing.

Machine-learning models, particularly anomaly-detection approaches trained on healthy-state data, can identify subtle correlated changes across dozens of signals that no fixed threshold captures. They adapt to operating regime through feature normalisation or regime clustering. But they require substantial clean labelled data to train, they degrade if the process changes without retraining, and their outputs (“anomaly score: 0.83”) are harder for a maintenance technician to act on than a clear rule (“bearing 3 surface defect frequency elevated 40% above baseline”).

PwC’s Predictive Maintenance 4.0 report and Deloitte’s research on predictive technologies for asset maintenance both identify that the organisations achieving the strongest returns combine rule-based alerting for known failure modes with model-based anomaly detection for unknown patterns, using domain engineers to translate model outputs into actionable maintenance guidance rather than expecting frontline technicians to interpret raw scores.

$1.4T
Annual unplanned downtime cost to Global 500, up 62% since 2019 (Siemens)
25–30%
Estimated reduction in maintenance costs achievable with mature PdM (Deloitte, PwC)
70%
Share of equipment failures that are random or age-independent, undermining calendar-based schedules (IBM)

Why PdM programmes stall

The gap between PdM pilot and PdM at scale is wide, and the reasons it persists are not primarily technical. Most organisations that have tried and stalled report the same cluster of failure modes.

Alert fatigue and false positives

A PdM system tuned too sensitively to avoid missing failures generates large numbers of alerts that do not lead to real failures. When technicians learn through experience that most alerts are spurious, they begin ignoring them, including the real ones. The trust destruction is rapid and recovery is slow. Threshold calibration and model tuning require dedicated time from both data engineers and domain maintenance experts, and that resource is rarely budgeted in the implementation plan.

Data silos and integration gaps

PdM platforms produce signals. Those signals need to reach a computerised maintenance management system (CMMS) to become work orders, and work orders need to carry the right procedures, the right parts list, and the right approval routing. In many plants, the sensor data lives in a historian, the alerts live in the PdM software dashboard, the CMMS sits in a separate enterprise system, and the maintenance procedures are in a shared drive or a paper folder. No single integration joins all four. The result is that an alert creates manual transcription work rather than automatic workflow.

No closed loop to the work order

Closely related to the silo problem: even where alerts do reach CMMS, the outcome of the maintenance task often does not flow back to improve the model. Whether the bearing was actually defective when the technician arrived, what the defect mode was, and what the root cause turned out to be are precisely the labelled data that would improve predictions over time. Without that feedback loop, the model stays at pilot-era accuracy indefinitely.

ROI that never materialises past the pilot

Pilots are usually run on the most instrumented, most critical assets where failure consequences are already well understood. The business case looks strong. When the programme is extended to a broader asset population, sensor density drops, failure history is sparse, domain knowledge is less concentrated, and the per-asset economics deteriorate. Organisations that base the full programme business case on pilot results often find the numbers do not hold at scale. IBM’s overview of the field notes that most well-designed programmes reduce maintenance costs by 10–25% and equipment breakdowns by a similar margin, but those figures assume full integration, not a dashboard running alongside existing processes.

Staffing and skills gaps

PdM sits at the intersection of reliability engineering, data science, and frontline maintenance, and few organisations have all three capabilities in-house at adequate depth. The reliability engineer who understands failure modes may not be confident in feature engineering; the data scientist who can build an anomaly model may not know what a gear mesh frequency looks like in a spectrum; the technician who acts on the alert may not have the diagnostic training to interpret it correctly. This skills triangle is a structural constraint, not a training problem that resolves quickly.

What separates programmes that scale

Organisations that move beyond pilot share several practices that differ from those that do not.

  • Starting with known failure modes. Successful programmes begin with asset classes where failure physics is understood, vibration signatures for bearing defects, motor current signatures for rotor bar cracks, thermal patterns for electrical faults. Model-based anomaly detection is added later, once the data infrastructure and alert culture are established on firmer ground.
  • Bidirectional CMMS integration. Alerts automatically generate work orders with the correct task list, required parts, and safety procedures. Work order completion flows outcomes back to the PdM platform so false positive rates can be tracked and models retrained. The loop is closed in both directions from day one.
  • Alert tiering and escalation logic. Rather than a binary alert/no-alert model, mature programmes use tiered severity levels: informational trends that go into a weekly review, advisory alerts that appear on a planner’s queue, and immediate alerts that open a work order automatically. This structure protects technicians from noise while ensuring critical signals are not buried.
  • Dedicated reliability engineer ownership. Scaling programmes almost universally have a named reliability engineer responsible for model accuracy, alert calibration, and outcome tracking, not as an add-on to their existing role but as a primary accountability. Without that ownership, alert drift and model decay are inevitable.
  • Phased asset rollout tied to criticality. Rather than instrumenting everything at once, programmes that scale prioritise by criticality tier, instrument the highest-consequence assets first, prove the economics, and use that proof to fund the next tier. This avoids the false-economy trap of broad shallow coverage.

The last mile: getting the signal to the technician

Even a well-calibrated PdM programme with clean CMMS integration still faces the last-mile problem: a maintenance technician standing in front of a machine with a work order that says “bearing anomaly detected, inspect and reseat if necessary” still needs to know which bearing, what the anomaly signature looked like, what normal looks like for comparison, and what the correct inspection and reseating procedure is for that specific component on that specific asset variant.

This is where connected worker technology intersects with PdM. Digital work instructions that carry the sensor trend, annotated with the expected failure mode and the step-by-step procedure, give the technician everything needed to execute correctly and capture evidence. Without that last-mile delivery, the system produces a work order that is executed on institutional memory and verbal hand-off, which reintroduces the variability and documentation gaps that the PdM investment was meant to eliminate.

The failure to solve the last mile is one reason PdM programmes produce measurable results in control rooms, visible on KPI dashboards, without producing a corresponding change in mean time to repair or first-time fix rates. The signal arrived; the execution was not upgraded to match.

Predictive maintenance is neither straightforward nor automatic, and framing it that way has probably set back more programmes than it has helped. The realistic path is a systematic reduction in reactive failure through progressively better instrumentation, integration, and execution quality, with each stage funded by the savings of the last. The organisations achieving the largest gains are not the ones who bought the most sophisticated analytics platform. They are the ones who closed the loop from signal to work order to outcome, built the organisational skills to maintain that loop, and resisted the temptation to declare victory after a successful pilot.

What is the difference between condition-based and predictive maintenance?

Condition-based maintenance (CBM) involves taking periodic measurements (vibration, oil samples, thermal readings) and acting when those readings exceed acceptance limits. It is still discrete and human-driven. Predictive maintenance uses continuous sensor streams and automated analysis to detect developing failures earlier and with less reliance on scheduled human data collection. In practice the terms overlap: a mature CBM programme with near-continuous monitoring is functionally predictive, while a PdM system that only checks data weekly is operationally closer to CBM.

How long does it typically take to see ROI from a PdM programme?

Pilot results on high-criticality assets often materialise within six to twelve months, particularly where the programme prevents even one major unplanned failure. Full-programme ROI, across a broader asset population with proper CMMS integration, typically requires eighteen to thirty-six months to stabilise. Programmes that report no return within two years usually have an integration gap (alerts not reaching work orders), a calibration problem (excessive false positives eroding trust), or a skills constraint that prevents model maintenance.

Which industries see the strongest gains from predictive maintenance?

Capital-intensive industries with high unplanned-downtime costs benefit most: oil and gas, continuous-process chemicals, power generation, automotive assembly, and mining. These sectors combine high failure consequence, well-characterised failure physics for rotating equipment, and large asset populations over which the cost of instrumentation is spread. Discrete manufacturing with shorter machine cycles and lower failure costs per event typically sees smaller absolute savings, though relative maintenance cost reductions can still be significant.

Do we need to replace existing sensors and historians to start?

Not necessarily. Many plants already have more sensor data than they use. The first step is usually to audit what signals are being collected, whether they are reaching any analytics layer, and whether that analytics layer is connected to the CMMS. Purpose-built vibration or oil analysis sensors may be added for critical assets, but the more common barrier is not sensor coverage, it is the absence of integration between the data that already exists and the work order system where action happens.

See a connected worker platform in action

Treedis turns your site into a digital twin, with every procedure, work order, and live reading pinned to the asset it belongs to.

Book a demo

We use cookies to improve your experience and analyse site traffic. Cookie policy.

DeclineAllow