What RCM Actually Is — And What It Isn’t
Reliability Centered Maintenance is a structured process for determining what must be done to ensure a physical asset continues to do what its users want it to do in its current operating context. That definition comes from John Moubray’s original work, and it still holds. RCM answers a deceptively simple question: what’s the right maintenance task for each failure mode of each piece of equipment?
What RCM is not: a software product, a maintenance management system, or an academic exercise. Too many plants have invested months in RCM analyses that produced three-ring binders nobody opened again. That’s an implementation failure, not a methodology failure.
The Seven Questions of RCM
Every RCM analysis answers seven questions for each asset. Work through these in order — skipping ahead leads to bad maintenance decisions.
- What are the functions? — Not what the equipment IS, but what it DOES in its operating context. A pump doesn’t just “pump.” It transfers 500 GPM of cooling water at 45 psig to the heat exchanger bank. Be specific. Include primary functions and secondary functions like containment, structural support, and environmental compliance.
- What are the functional failures? — How can it fail to perform each function? Total loss of function is obvious. Partial function loss (pump delivers 350 GPM instead of 500 GPM) is equally important and often overlooked.
- What are the failure modes? — What specific events cause each functional failure? Impeller wear, mechanical seal failure, coupling failure, bearing seizure — these are failure modes. Be specific enough to select a maintenance task, but don’t go so deep that you’re analyzing at the individual bolt level.
- What are the failure effects? — What happens when each failure mode occurs? Describe the evidence, the safety impact, the environmental impact, and the operational impact. Include what the operator sees, hears, and experiences. This information drives the consequence evaluation.
- What are the consequences? — RCM classifies consequences into four categories: hidden failures (safety/environmental protective devices), safety/environmental, operational (production impact), and non-operational (repair cost only). The consequence category determines which types of maintenance tasks are acceptable.
- What proactive tasks are applicable and effective? — For each failure mode, evaluate condition-based tasks (predictive maintenance), scheduled restoration, and scheduled discard. A task is applicable if it can detect or prevent the failure mode. It’s effective if the cost of the task is less than the cost of the failure consequence.
- What should be done if no proactive task is found? — Run-to-failure is a valid strategy for non-critical failure modes with low consequences. For hidden failures, failure-finding tasks (periodic functional tests) become the default. For safety consequences with no applicable proactive task, redesign is the answer.
Streamlining the Process: Where Teams Get Stuck
Classical RCM analysis per SAE JA1011/JA1012 is rigorous and thorough. It can also consume enormous amounts of time if you let it. A full RCM analysis on a complex system like a boiler or turbine generator can take weeks of facilitated sessions. Most plants can’t sustain that level of effort across their entire asset base.
Focus on Critical Assets First
Don’t try to analyze everything. Rank your assets by criticality — production impact, safety risk, environmental risk, and repair cost. The top 10-20% of your asset base drives 80-90% of your maintenance burden and failure risk. Start there.
Use Existing Data
Your CMMS work order history, operator logs, and reliability data contain a wealth of failure mode information. Before convening an RCM team, mine your data for the actual failure modes your plant experiences. This eliminates the theoretical failures that waste analysis time and focuses the team on real problems.
Right-Size the Analysis Depth
Not every asset needs a full classical RCM analysis. SAE JA1012 itself acknowledges that streamlined approaches are acceptable when the analyst understands the trade-offs. For critical, complex systems with high consequences of failure — full RCM. For moderately critical systems with well-understood failure patterns — a streamlined analysis using templates from similar equipment is sufficient. For low-criticality assets — a simple decision based on consequence assessment works fine.
Keep the Team Small and Focused
The RCM review team should include an operations representative, a maintenance technician with hands-on experience on the equipment, an engineer, and a facilitator trained in RCM methodology. Six people maximum. Larger groups slow the process dramatically without improving the analysis quality.
From Analysis to Action: Implementing RCM Results
The analysis is only valuable if it changes what you actually do. Implementation requires translating RCM decisions into actionable maintenance tasks in your CMMS.
Condition-Based Tasks
RCM frequently identifies condition-based monitoring (predictive maintenance) as the preferred strategy for failure modes with detectable degradation patterns. These tasks need to specify: what condition to monitor, what technology to use, what the inspection interval should be, and what the alert criteria are.
The inspection interval should be less than half the P-F interval — the time between when a failure is detectable and when it becomes functional failure. If vibration analysis can detect a bearing defect six months before failure, your monitoring interval should be three months or less.
Time-Based Tasks
For failure modes with age-related patterns (rubber degradation, corrosion, fatigue), scheduled restoration or replacement at fixed intervals makes sense. The interval should be set based on the onset of the wear-out failure pattern, adjusted for operating conditions and consequence severity.
Failure-Finding Tasks
Hidden failures — like a relief valve that sits dormant until needed — require periodic functional testing. The test interval depends on the probability of the hidden failure coinciding with the demand for the protective function. For critical safety systems, test intervals of 3-6 months are common. Test the function, not just the component.
Run-to-Failure
This is a deliberate decision, not neglect. When the cost of preventive or predictive maintenance exceeds the cost of failure, run-to-failure is the correct strategy. Document these decisions so future maintenance planners don’t add unnecessary tasks. Stock appropriate spares to support rapid repair when failure occurs.
Measuring RCM Success
Track these metrics before and after RCM implementation:
- Unplanned downtime — Should decrease by 25-50% on analyzed assets within 18 months
- Emergency work orders as percentage of total — Target less than 10%. Pre-RCM plants often run 30-50% emergency work.
- PM task count — Counter-intuitively, RCM often reduces the number of PM tasks. Plants frequently discover they’re performing unnecessary PMs on failure modes that don’t justify scheduled intervention.
- Maintenance cost per unit of production — Should trend downward as reactive work decreases and planned work increases efficiency.
RCM isn’t a project with a completion date. It’s a living process. As operating conditions change, new failure modes emerge, and equipment ages, your maintenance strategies need to evolve. Schedule annual reviews of RCM decisions for critical assets. Update the analysis when significant failures occur that weren’t anticipated. The plants that sustain RCM discipline over years see compounding benefits that dwarf the initial implementation effort.