A critical centrifugal pump fails during production. The maintenance team replaces the mechanical seal, checks for obvious leakage, and returns the pump to service before the outage becomes a larger event. Later, the same seal fails again, often after the operating conditions, alignment, lubrication practice, and installation quality have remained largely unchanged.
That repair restored the function, but it didn't necessarily restore reliability. Root cause analysis FMEA connects the immediate event to the broader maintenance strategy. Root cause analysis investigates why the failure happened, while Failure Mode and Effects Analysis, or FMEA, ensures the confirmed cause changes how the plant prevents, detects, and prioritizes that failure in the future.
Table of Contents
- The Hidden Cost of Recurring Equipment Failures
- Comparing Root Cause Analysis and FMEA Methodologies
- Integrating FMEA into Root Cause Investigations
- Calculating Risk Priority Numbers for Rotating Equipment
- Overcoming Common FMEA Blind Spots in Plant Environments
- Measuring Reliability Program Success and Metrics
- Next Steps for Your Reliability Strategy
The Hidden Cost of Recurring Equipment Failures
A seal replacement can be technically correct and still be an incomplete maintenance action. If a pump shaft is misaligned, the seal may be absorbing vibration that originates elsewhere. If the fluid contains contamination, the seal faces may wear prematurely. If the pump operates away from its intended duty point, hydraulic forces can create conditions that a new seal won't solve.
The visible failure is only one layer of the problem. The seal leak is the failure effect or symptom. The underlying cause may be misalignment, pipe strain, lubrication loss, contamination, excessive vibration, or an unsuitable operating condition. Treating the effect without verifying the cause leaves the next failure mode active.
What root cause analysis contributes
Root cause analysis, or RCA, is a backward-looking investigation. The team starts with an actual event, gathers evidence, reconstructs the sequence, and tests possible physical, human, and process causes. A sound investigation doesn't stop at “the seal failed” or “the bearing was worn.” It asks what conditions produced that damage and what evidence supports the conclusion.
For the pump, useful evidence may include:
- Physical condition: Seal-face damage, shaft scoring, bearing condition, coupling wear, and deposits on wetted components.
- Operating history: Suction conditions, discharge pressure, flow changes, starts and stops, and process upsets.
- Maintenance records: Installation method, alignment results, lubrication work, previous repairs, and parts used.
- Condition data: Vibration spectra, temperature trends, oil condition, ultrasound observations, and motor current behavior.
The investigation should distinguish a confirmed cause from a plausible explanation. “The pump was old” isn't a cause that maintenance can control. “Soft foot caused coupling misalignment, which increased radial movement at the seal” is a testable causal statement.
What FMEA contributes
FMEA is a forward-looking control method. The team asks how an asset or process could fail, what effect each failure would have, what could cause it, and which controls can prevent or detect it. FMEA then ranks the risk so limited engineering and maintenance resources go toward the most consequential failure mechanisms.
The method has a long reliability history. The U.S. military described FMEA in MIL-P-1629 in 1949, NASA adopted it for the Apollo program in 1963, and its use expanded into aviation, nuclear, food, and automotive applications. Later standardization included DIN 25448 in 1980, IEC 60812 in 1985 and 2001, and the AIAG/Ford/GM FMEA Reference Manual in 1993. These milestones are documented in the history of FMEA and failure analysis.
A reactive repair keeps production moving. An integrated RCA and FMEA process changes the asset strategy, the inspection route, the work instructions, and the risk record so the same failure isn't repeatedly rediscovered during an audit, shutdown, or emergency callout.
Comparing Root Cause Analysis and FMEA Methodologies
RCA and FMEA answer different questions, and confusing their roles creates weak reliability decisions. RCA asks, “Why did this failure happen?” FMEA asks, “How could this asset fail, and what should be done before it does?”
The distinction matters on a gearbox, pump, compressor, or turbine train. An RCA may establish that bearing damage resulted from lubricant contamination. The FMEA must then determine whether contamination is already represented as a cause, whether the existing detection control can find it early, and whether the maintenance plan should change.
RCA vs FMEA Comparison Matrix
| Dimension | Root Cause Analysis (RCA) | Failure Mode and Effects Analysis (FMEA) |
|---|---|---|
| Primary direction | Backward-looking, based on an actual event | Forward-looking, based on potential failure modes |
| Starting point | A breakdown, defect, unsafe condition, or performance loss | An asset, subsystem, process, or operating function |
| Main question | Why did the event occur? | How could failure occur, and what would the effect be? |
| Evidence | Inspection findings, operating data, work orders, interviews, and tests | Asset knowledge, drawings, failure history, controls, and operating context |
| Main output | Verified causal logic and corrective action | Prioritized risks, preventive controls, detection controls, and ownership |
| Timing | After a failure or significant deviation | Before failure, during design or planning, and after new evidence arrives |
| Equipment example | Explains why a pump seal leaked | Identifies seal leakage, bearing wear, or coupling failure as risks |
| Update requirement | Closes the investigation and assigns corrective actions | Revises causes, ratings, controls, and action priorities |
A plant that performs RCA without updating FMEA gains a local explanation but loses organizational learning. A plant that maintains FMEA without feeding it field evidence ends up with a theoretical document that may not reflect actual duty cycles, workmanship, contamination, or degradation patterns.
The RCA and failure analysis workflow should therefore treat the two methods as a closed loop. The investigation confirms what happened. The FMEA records what the organization now knows and changes the controls that govern future work.
Choosing the right method
Use RCA when an asset has failed, a defect has escaped, a protective device has acted unexpectedly, or a recurring problem demands evidence rather than opinion. The investigation should be proportional to the consequence and complexity of the event.
Use FMEA when commissioning a new process, reviewing a critical asset, changing an operating condition, revising a maintenance plan, or incorporating lessons from an RCA. FMEA also provides a useful structure before a shutdown because it forces the team to identify likely failure modes and define what should be checked.
The methods aren't competing alternatives. RCA supplies verified field learning, while FMEA turns that learning into repeatable prevention and detection.
Integrating FMEA into Root Cause Investigations
The handoff from RCA to FMEA should be a controlled workflow, not an informal note in a closeout report. A failure investigation is complete only when the confirmed cause has been translated into an updated risk record, a specific control, and an accountable action.

Start with the actual failure mode
First, define the failure mode precisely. “Pump problem” is too broad. “Mechanical seal leakage at the drive-end pump” identifies the functional loss and gives the team a usable starting point.
The team should record the effect separately from the failure mode and cause:
- Function: Transfer process fluid at the required operating condition.
- Failure mode: Mechanical seal leaks.
- Effect: Fluid loss, environmental exposure, reduced availability, or forced shutdown.
- Cause: Shaft misalignment, pipe strain, contamination, thermal distortion, or installation error.
This separation prevents the FMEA from collapsing several layers into a single vague row. It also makes the later control decision more precise.
Trace the cause with a suitable RCA method
For a straightforward failure with a clear chain, the 5 Whys method provides an iterative path from symptom toward cause. It was developed inside the Toyota Production System by Taiichi Ohno and later codified in his 1988 book, reflecting a shift from correcting symptoms to eliminating underlying causes. The method should continue until the team reaches a condition that can be verified and controlled, not just because a predetermined number of questions has been asked.
For a more complex system, fault tree analysis, or FTA, maps combinations of events that can produce a top-level failure. It is a quantitative causal diagram and requires known component failure-rate data, making it more suitable for complex equipment such as turbine trains or critical process skids. The distinction between these methods is described in the NASA comparison of root cause analysis methods.
An Apollo investigation or a 5-Why investigation should produce a causal statement that can be placed into the FMEA cause column. The word “Apollo” here refers to a structured causal investigation approach, not to a substitute for physical evidence. If the pump seal failed because alignment shifted after thermal growth and the existing inspection method couldn't detect it under operating conditions, both facts belong in the FMEA.
Re-score the risk and revise the controls
Once the cause is confirmed, the FMEA team should review the failure mode's severity, occurrence, and detection ratings independently. The team shouldn't lower the occurrence rating merely because a new inspection was added, and it shouldn't let a high severity rating inflate the detection score.
The updated row should identify:
- Prevention control: Correct alignment procedure, pipe-support verification, thermal-growth review, or installation quality check.
- Detection control: Operating vibration analysis, coupling inspection, seal leakage monitoring, or a targeted route observation.
- Action owner: The person responsible for implementation, not merely the department.
- Verification method: The evidence required to show that the action works.
- Review trigger: The condition that requires another assessment, such as a process change, repeat event, or control failure.
This is the practical link between a scored risk item and a field investigation. The FMEA approach for manufacturing reliability should result in a changed work instruction, route, design decision, spare-parts requirement, or operating limit, not only a revised spreadsheet cell.
Close the loop in the maintenance system
The FMEA update should connect to the CMMS or equivalent work-management process. If the new control is vibration analysis, the route needs an asset, measurement point, frequency, alarm logic, and reaction plan. If the action is improved seal installation, the job plan needs the correct tolerances, inspection steps, and acceptance evidence.
Practical rule: A root cause isn't closed until the plant can show what changed in the way the asset is operated, maintained, inspected, or designed.
Calculating Risk Priority Numbers for Rotating Equipment
The Risk Priority Number, or RPN, is calculated as severity multiplied by occurrence multiplied by detection. It helps a team sort failure modes and focus limited resources on the risks that deserve attention first. The number is useful only when the underlying ratings are disciplined and the failure logic is clear.
Consider a process pump and its driver gearbox. The analysis should not use “pump failure” as one generic entry. A useful FMEA separates the failure mode, effect, and cause.
Separate the failure layers
For gearbox bearing wear, the structure might look like this:
- Failure mode: Bearing wear or spalling.
- Effect: Increased vibration, gear-mesh disturbance, heat generation, and eventual loss of torque transmission.
- Potential causes: Lubrication loss, contamination, misalignment, overload, or electrical stress transmitted through the drive system.
- Prevention controls: Correct lubrication selection, sealing improvements, alignment verification, load review, and installation quality.
- Detection controls: Vibration analysis, oil analysis, temperature monitoring, and inspection of abnormal operating conditions.
For pump seal leakage, the structure changes:
- Failure mode: Mechanical seal leakage.
- Effect: Product loss, contamination risk, reduced pump availability, or forced shutdown.
- Potential causes: Shaft movement, misalignment, pipe strain, seal-face damage, dry running, or unsuitable process conditions.
- Prevention controls: Correct seal selection, installation checks, flush-plan verification, and operating procedure review.
- Detection controls: Leakage observation, vibration monitoring, pressure and flow review, and condition checks during rounds.
This separation matters because a component-level action may miss the actual mechanism. Replacing a bearing doesn't correct lubrication contamination. Replacing a seal doesn't correct shaft movement.
Score the three factors independently
Severity describes the consequence if the failure occurs. Occurrence describes the likelihood or frequency of the cause under the relevant operating conditions. Detection describes how likely the existing controls are to identify the developing failure before the effect becomes serious.
The team must score these factors separately and independently. A severe consequence shouldn't automatically receive a high occurrence score, and a frequent historical event shouldn't automatically receive a poor detection score. Independent scoring improves consistency when comparing gearbox bearing wear with pump seal leakage.
A high RPN can result from different combinations of severity, occurrence, and detection. That means a lower total number shouldn't automatically make a failure safe to ignore. A failure with serious consequences may warrant an action even when its occurrence is low, especially if detection is weak.
The RPN method for reliability prioritization is most useful when the team documents the rating rationale. “Detection is low because the current operator round checks leakage but doesn't measure vibration” is more actionable than an unexplained score.
Turn scoring into maintenance decisions
The score should lead to a decision about prevention, detection, or both. For bearing wear, oil analysis may reveal contamination or wear debris, while vibration analysis can identify developing mechanical deterioration. For seal leakage, visual checks may detect an active leak, but vibration and alignment verification can address the mechanism before leakage occurs.
The FMEA should also support preventive maintenance interval optimization, spare-parts prioritization, and condition-monitoring decisions for critical rotating equipment. It shouldn't prescribe a calendar task merely because a component is important. The chosen task must correspond to a recognizable degradation mechanism and a useful intervention point.
Overcoming Common FMEA Blind Spots in Plant Environments
A conference-room FMEA can look complete while missing the failures that matter on the plant floor. Broad brainstorming often produces familiar entries such as “bearing failure,” “seal failure,” or “pump stops.” Those labels identify component outcomes, but they don't reveal which service condition, work practice, or interaction causes the damage at a particular site.
The stronger approach compares similar assets that operate differently. Two centrifugal pumps may share a design but face different suction conditions, fluid contamination, duty cycles, operator practices, alignment quality, and maintenance histories.

Replace abstract prompts with evidence
A useful session starts with actual failure history, work orders, inspection findings, and operating context. The facilitator can compare a pump that repeatedly loses seals with a similar pump that remains stable, then ask what differs between the two assets.
Relevant differences may include:
- Duty cycle: Continuous operation, frequent starts, standby service, or process cycling.
- Fluid condition: Solids, viscosity changes, temperature variation, or chemical compatibility.
- Installation quality: Base condition, soft foot, coupling alignment, pipe strain, and support integrity.
- Maintenance execution: Lubricant quantity, seal handling, torque practice, cleanliness, and post-work testing.
- Monitoring coverage: Whether the route detects the degradation mechanism or only observes the final symptom.
This method exposes failure interactions that a generic worksheet misses. For example, a seal may fail after a bearing begins to deteriorate, not because the seal itself was defective. A gearbox may experience accelerated wear when lubrication contamination combines with misalignment and variable loading.
Control bias in the room
FMEA quality depends on who participates and what evidence they bring. A design specialist may understand the equipment architecture but miss an installation shortcut. A maintenance planner may know the job plan but not the process upset that repeatedly overloads the asset. An operator may recognize the early sound or pressure change that never appears in the work-order description.
The facilitator should challenge unsupported certainty and distinguish known facts from assumptions. The contributing-factor analysis approach helps the team examine the conditions around the event without reducing the investigation to individual blame.
The evidence-based approach aligns with guidance that stronger FMEA results come from comparing similar failed assets, examining differences in service conditions, and using specific prior failure history rather than brainstorming in the abstract. That practical blind spot is discussed in guidance on improving FMEA interview and facilitation quality.
A generic FMEA describes what equipment can do. A site-specific FMEA describes what this equipment does under this plant's conditions.
Measuring Reliability Program Success and Metrics
An updated FMEA has value only when it changes field behavior and asset performance. The measurement system should therefore connect the investigation to the work that follows it, then connect that work to operational results.
Unplanned downtime is the most direct measure for a recurring pump, gearbox, or compressor failure. The team should track whether the targeted failure mode returns, how long the asset remains unavailable, and whether the corrective action prevents recurrence under comparable service conditions. A single successful repair doesn't prove that the cause has been controlled.
Use metrics that expose the mechanism
Plant leaders often need a concise view, while reliability engineers need diagnostic detail. A practical scorecard can include:
- Repeat failure frequency: Whether the same failure mode returns after the RCA action.
- Unplanned downtime: Lost operating time associated with the targeted asset or failure mechanism.
- Mean time between failures: The interval between relevant functional failures, interpreted with care when operating conditions change.
- Mean time to repair: The time required to restore service, including access, diagnosis, parts, and testing.
- Overall Equipment Effectiveness: The combined view of availability, performance, and quality, used alongside the asset-level evidence rather than as a standalone explanation.
- Action closure quality: Whether actions have owners, due dates, verification evidence, and an updated FMEA row.
Definitions and relationships among MTBF, MTTR, and OEE are summarized in this reliability metrics guide.
Measure detection, not just failure
A predictive maintenance route can appear active while failing to identify the mechanism that matters. If vibration analysis is assigned to a bearing failure but measurements aren't taken at the correct points or under comparable operating conditions, the route may create activity without useful detection. Oil analysis has the same limitation when samples are contaminated, delayed, or disconnected from an action threshold.
Thermography can help identify abnormal heating in electrical and mechanical systems. Motor current signature analysis can reveal electrical or mechanical behavior through the motor's current signal. These techniques should enter the FMEA as specific detection controls with a defined reaction, not as a list of technologies.
The team should ask whether the control detects degradation early enough to act. If it doesn't, the response may require a different measurement, a design change, an operating control, or a more direct preventive action.
Use FMEA to guide investment
The revised risk record can support spare-parts decisions, maintenance interval changes, and capital planning. A critical gearbox with weak detection and a known lubrication-related cause may justify improved sealing, filtration, oil sampling, or a design review. A pump with a recurring alignment-related failure may need installation improvements rather than more frequent seal replacement.
The important result isn't a favorable metric by itself. The result is a traceable chain from failure evidence to risk ranking, from risk ranking to control, and from control to verified plant performance.
Next Steps for Your Reliability Strategy
Recurring failures continue when the organization treats each repair as an isolated event. The maintenance team replaces the damaged component, closes the work order, and moves on, while the FMEA remains unchanged and the next technician inherits the same exposure.
A stronger strategy makes the handoff mandatory. Every significant RCA should answer four operational questions:
- What failed, and what evidence confirms the failure mode?
- What physical, process, or human conditions caused it?
- Where is that cause recorded in the relevant FMEA?
- What prevention, detection, or design control now changes the risk?
The process also needs a clear threshold for investigation. Not every minor defect requires a complex causal study, but repeat failures, consequential breakdowns, unexplained condition changes, and failures that defeat existing controls deserve more than component replacement.
Plant leaders should review whether critical assets have current failure logic, whether condition-monitoring routes detect actual degradation mechanisms, and whether completed RCA actions change job plans and control plans. If the answer is no, the gap isn't only documentation. It represents a risk that remains active in the field.
A free reliability assessment can expose that gap by examining critical rotating assets, recurring work orders, FMEA quality, monitoring coverage, and maintenance decision logic. The assessment should produce practical priorities, not a generic list of improvement themes.
Forge Reliability offers a free reliability assessment that connects recurring failure investigations with FMEA updates, condition-monitoring routes, and maintenance actions for critical equipment. Visit Forge Reliability to evaluate hidden downtime risks and build a focused plan for preventing the next pump, gearbox, compressor, or motor failure.