Home / Blog / How to Perform Root Cause Analysis: Step-by-Step Guide 2026
Reliability Engineering Insights

How to Perform Root Cause Analysis: Step-by-Step Guide 2026

16 min read ·
How to Perform Root Cause Analysis: Step-by-Step Guide 2026

A centrifugal pump seizes again on second shift. Operations reports rising motor amps before the trip, maintenance finds heat at the inboard bearing housing, and production wants the spare installed before the line misses another batch. By the next morning, the team has three opinions and no verified answer. One blames lubrication. Another points to suction issues. A third says the operator ran it too far left on the curve.

That scene is common in manufacturing, chemical processing, food and beverage, water treatment, and any plant with critical rotating equipment. The immediate repair gets the asset running, but the same failure mode comes back because the investigation never moved past educated guesses. The plant keeps paying for the same lesson.

A disciplined approach to how to perform root cause analysis changes that pattern. It gives the team a method for selecting the right analysis tool, gathering evidence that holds up under review, building a defensible cause path, and assigning corrective actions that successfully prevent recurrence. It also forces one thing that many guides skip: tying every step to CMMS records so the learning survives beyond the meeting room.

For teams that also need a structured workflow to orchestrate incident resolution, the same principle applies. Incidents move faster when event history, actions, and ownership are visible instead of scattered across emails and shift notes. The same discipline that improves response quality also improves RCA quality.

A strong investigation also distinguishes the event itself from the conditions that allowed it. Teams that need a sharper way to frame those distinctions can use this guide to define contributing factors before they start assigning blame or rewriting procedures.

Table of Contents

Introduction to Structured Root Cause Analysis

A structured RCA starts with a simple idea. The failure has to be explained by evidence, not by the loudest opinion in the room. In a pump seizure, that means the team doesn't stop at “bearing failure” because a failed bearing is an observation, not a root cause.

The better question is what made the bearing fail under those operating conditions. Was the grease incompatible. Did shaft misalignment increase load. Did low flow create heat and internal recirculation. Did the seal flush fail and contaminate the housing. Every one of those paths requires proof.

What good RCA looks like

A useful RCA does three things at the same time:

  • Defines the problem clearly: The team writes the failure as a specific event, such as pump P-204 tripped on overload and seized during transfer service at a known date and operating state.
  • Builds a reliable timeline: The sequence of alarms, operator actions, PM history, parts changes, and process conditions becomes the backbone of the investigation.
  • Produces corrective actions with owners: The output isn't a generic note to “watch lubrication more closely.” It is a controlled action assigned to someone with the authority to complete it.

Practical rule: If the team can't show the sequence of what changed before the failure, it isn't ready to argue about causes.

Why ad hoc investigations fail

Many plants still treat RCA as an after-action meeting. The team gathers around a whiteboard, asks “why” a few times, writes down likely causes, and moves on. That approach feels efficient, but it usually preserves the same weak assumptions that caused the failure to return.

For reliability engineers and maintenance managers, structured RCA is less about paperwork and more about repeatability. A documented method lets the next investigator review the event, challenge the logic, and connect the findings to maintenance strategy. That matters when the asset is a process pump, an air compressor, a gearbox on a finishing line, or a VFD-driven fan in a dusty environment. In each case, the failure mechanism is different, but the discipline is the same.

Choosing the Right RCA Method

Not every failure needs the same level of rigor. A loose terminal in a motor starter and a recurring compressor shutdown don't deserve identical workflows. The method has to match the consequence, the complexity, and the quality of available evidence.

The biggest mistake is treating all RCA methods as equivalent. They aren't.

Where simple methods fit

5 Whys works best when the failure path is short, the evidence is direct, and the event doesn't involve many interacting conditions. It can help a team trace a missed lubrication route, a wrong part installation, or a delayed response to an alarm.

Fishbone diagrams help when the team needs to organize possible causes by category, such as machine, method, material, people, or environment. They are useful early in the process because they widen the search and keep the team from locking onto one theory too soon.

Those tools are valuable, but they are not enough on their own for high-consequence or technically messy failures.

A key limitation is bias control. According to William Hovey's discussion of why RCA efforts fail, approximately 70% of investigations lack independent management agreement and structured verification of all hypothesized causes, which pushes teams toward thought-provoking tools that don't filter bias. That matters in plants where a politically convenient answer can arrive faster than a technically correct one.

When to use structured logic

For complex equipment failures, Fault Tree Analysis and Apollo-style cause mapping are stronger choices because they force explicit cause-and-effect logic.

Use them when the event has these characteristics:

  • Multiple interacting causes: A pump seizure linked to low NPSH, seal leakage, contamination, and delayed operator response.
  • Conflicting evidence: Vibration trends point one way, teardown evidence points another, and process historians show a third issue.
  • Repeat failures: The plant has already tried several fixes and the same mode keeps returning.
  • Cross-functional ownership: Operations, maintenance, engineering, and planning all control part of the solution.

A structured method also pairs naturally with preventive tools. Teams that already use FMEA for manufacturing often adapt more quickly to rigorous RCA because they are used to thinking in failure modes, effects, and control gaps.

A brainstorming tool can surface possibilities. It can't prove causality by itself.

A practical selection guide

For plant-floor decisions, this rule set works well:

  1. Use 5 Whys for single-path failures with low complexity and clear evidence.
  2. Use Fishbone when the team needs broad hypothesis generation before narrowing.
  3. Use causal mapping or fault trees when evidence is mixed, recurrence is high, or the asset is operationally critical.
  4. Escalate the method if the first pass produces opinions instead of verifiable cause links.

A pump that seized once after obvious contamination may not need an elaborate tree. A pump that has seized three times across different shifts, after different rebuilds, absolutely does. Method selection is not about preference. It is about matching rigor to risk.

Planning Your RCA Investigation

A good RCA starts before the first meeting. Scope, authority, and team composition decide whether the investigation becomes a learning exercise or another argument between departments.

One widely used benchmark for quality is that a high-quality root cause analysis is defined by 18 measurable grading factors, including a complete timeline, verified causal relationships, and corrective actions that are specific, measurable, and assigned to authorized individuals, as outlined in TapRooT's guidance on good RCA. That standard is useful because it forces discipline early.

An infographic titled Planning Your RCA Investigation outlining five key steps for successful root cause analysis projects.

Set the boundaries first

The investigation scope has to answer three questions:

  • What failed: Name the asset and the top event precisely. “Pump bad” is useless. “Transfer pump P-204 seized during caustic circulation” is workable.
  • What impact matters: Production loss, safety exposure, quality risk, environmental release, or maintenance cost.
  • What time window applies: Include the hours or days before failure when process changes, repairs, or abnormal alarms started to accumulate.

A narrow scope can miss important contributors. A broad scope creates noise. For example, if a boiler feedwater pump trips after a recent seal replacement, the team should include the work order history, alignment records, and operating conditions before and after the repair, but it doesn't need to reopen every failure that asset has had in the last five years unless a pattern emerges.

Protect the investigation from politics

Management support matters, but so does management restraint. The team needs permission to follow evidence wherever it leads, including into planning errors, PM design flaws, spare parts issues, or operating practices. If leaders signal the answer in advance, the analysis is compromised before it begins.

The investigation lead should have authority to ask for records, interview personnel, and challenge assumptions without needing to defend every request.

Assign roles and operating rhythm

An effective RCA team usually needs a small core group with defined responsibilities:

  • RCA lead: Facilitates the logic, controls the evidence trail, and keeps the team from jumping to solutions.
  • Technical SME: Explains failure mechanisms, operating limits, and asset design details.
  • Data analyst or planner: Pulls historian data, CMMS records, inspection history, and parts usage.
  • Operations representative: Supplies real operating context, shift conditions, and procedural realities.
  • Action owner: Confirms whether proposed fixes are feasible, funded, and schedulable.

Meeting cadence also matters. Short working sessions with homework between them usually outperform one long workshop because people need time to collect records, inspect components, and test assumptions. Every session should end with evidence requests, not vague discussion points.

Collecting Evidence and Diagnosing Equipment Failures

The evidence phase decides whether the RCA will survive contact with reality. Many investigations often weaken here, especially on older assets that don't have continuous monitoring.

A 2024 study found that 42% of recurring equipment failures in U.S. manufacturing plants involve assets without predictive monitoring, yet fewer than 15% of RCA methods include protocols for data-deficient scenarios, according to the referenced industry discussion. For reliability teams working with legacy pumps, belt-driven fans, steam traps, and old conveyors, that gap is familiar.

Capture the failure story before it disappears

The first task is preserving evidence that degrades quickly. That includes process conditions, human observations, and physical component state.

For a recurring pump seizure, the collection sequence should include:

  • Operating state: Flow, suction and discharge pressure, fluid temperature, valve positions, tank level, and whether the pump was starting, stopping, or running steady.
  • Protection history: Motor trips, overload logs, vibration alarms, seal leak alarms, and temperature excursions.
  • Maintenance record: Last rebuild, bearing change, lubrication activity, alignment check, seal replacement, and spare parts source.
  • Physical condition: Shaft freedom, coupling condition, bearing discoloration, lubricant appearance, wear pattern, impeller damage, and evidence of rubbing or cavitation.
  • Human observations: Operator notes about noise, smell, heat, unstable flow, or process upset in the hours before failure.

A timeline should combine all of that into one sequence. That timeline belongs in the CMMS work order or linked investigation record, not in someone's notebook.

Handle data-deficient assets with field evidence

When sensors are sparse or noisy, the team still has options. The mistake is assuming no online data means no defensible RCA.

A practical protocol for data-deficient assets looks like this:

  1. Interview operators separately. Independent interviews reduce group memory bias and often reveal sequence details that get flattened in a team meeting.
  2. Inspect the failed parts before cleanup. Surface scoring, heat tint, lubricant contamination, and fracture location can narrow the failure path fast.
  3. Review route-based condition data. Monthly vibration or infrared routes may be the only trend available, but they often show whether the condition was worsening.
  4. Pull contextual process records. Batch changes, line blockages, flush interruptions, and suction source changes often explain why the machine saw abnormal duty.
  5. Compare with the healthy sister asset. If two pumps share service and only one fails, differences in alignment, pipe strain, base condition, or operator use can be revealing.

For teams building route-based monitoring on legacy machines, a practical starting point is understanding how to measure vibrations consistently enough that those readings can support later RCA work.

Sparse data doesn't remove the need for proof. It changes where the proof comes from.

Diagnostic Techniques by Equipment Type

Equipment Type Diagnostic Technique Key Data
Centrifugal pump Vibration, bearing inspection, suction and discharge review, seal and impeller teardown Bearing condition, hydraulic instability, alignment evidence, contamination signs
Compressor Motor current review, temperature trending, lubricant condition, valve and clearance inspection Load changes, thermal distress, lubrication state, process upsets
Gearbox Vibration route, oil debris inspection, backlash and tooth contact check Gear wear pattern, lubrication condition, mounting integrity
Electric motor Insulation checks, current signature review, bearing examination, coupling inspection Electrical imbalance, rotor or stator symptoms, bearing distress, misalignment clues
Heat exchanger or static vessel support equipment Process trend review, operator logs, inspection records, fouling evidence Differential temperature changes, blockage, pressure drop, cleaning history

Build an evidence package that can be tested

The strongest evidence packages mix direct and indirect proof. A direct proof item could be a bearing with clear contamination ingress. An indirect proof item could be a pattern of increased housing temperature after every washdown cycle. Together they make a stronger case than either one alone.

That matters in industries where process conditions shift often. In food plants, CIP cycles can change temperature and seal conditions. In chemical units, fluid properties may vary by batch. In mining and aggregate service, dust contamination and load swings can overwhelm a marginal design. The RCA has to capture the operating reality of that equipment in that plant, not a generic textbook version of the asset.

Building Causal Trees and Validating Root Causes

A causal tree turns scattered facts into logic. It starts with the top event and breaks it into the conditions that had to exist for the failure to occur. That structure is what keeps the team from stopping at symptoms.

Teams using only 5 Whys achieve a 45% success rate in root cause identification, while teams combining Fishbone with statistical correlation reach 78%. The same source reports that 52% of analyses fail to validate causes before action, doubling downtime recurrence, according to the Six Sigma RCA methodology summary. The lesson isn't that simple tools are useless. It is that unvalidated logic is expensive.

A diagram illustrating the process of building causal trees and performing root cause analysis using Fault Tree Analysis.

Start with a precise top event

The top event should describe the observable failure, not the conclusion. “Pump failed from poor maintenance” is already biased. “Pump P-204 motor tripped on overload and rotor could not be turned by hand” is a valid top event.

From there, the tree should expand at least three layers deep on technically significant failures. A practical structure might look like this:

  • Top event: Pump seizure during transfer service
  • Primary cause branch: Rotating assembly experienced abnormal friction
  • Secondary branch: Bearing distress
  • Tertiary branch: Lubricant contamination from failed seal barrier
  • Tertiary branch: Incorrect grease type introduced during PM
  • Secondary branch: Rotor contact
  • Tertiary branch: Pipe strain after maintenance
  • Tertiary branch: Thermal growth not accommodated
  • Primary cause branch: Pump operated outside stable hydraulic range
  • Secondary branch: Minimum flow protection ineffective
  • Tertiary branch: Control valve stuck
  • Tertiary branch: Bypass logic disabled

That level of breakdown forces the team to show what had to be true, not what sounds plausible.

For teams that want a more visual way to organize these linkages, cause-and-effect maps can help convert raw failure discussion into a reviewable logic structure.

Validate every branch with evidence

A causal node should only stay in the tree if the team can support it. Validation can take several forms:

  • Correlation: Process trend changes align with the onset of abnormal temperature or current.
  • Physical confirmation: Teardown findings match the predicted failure mechanism.
  • Comparison: Sister asset data shows the failed unit behaved differently under the same duty.
  • Pilot test: A controlled operating or maintenance change removes the suspected condition.
  • Exclusion: Evidence rules out a competing theory.

Field note: If two branches remain possible, the team shouldn't pick one by consensus. It should gather evidence until one branch weakens or both remain as conditional contributors.

A food processing example makes this concrete. A product transfer pump may show repeated bearing failures after washdown. The initial blame often lands on lubrication frequency. But a proper tree may reveal that high-pressure washdown entered through a degraded seal arrangement, contaminated the bearing cavity, and was made worse by a PM task that overgreased the housing. In that case, lubrication is part of the story, but not the full causal path.

Validation also protects the team from management reviews that ask the right hard questions. If the branch says “operator error,” the team needs evidence of what the operator saw, what the procedure required, whether the indication was clear, and whether the process conditions made the action reasonable. Root causes must survive scrutiny, not just workshop agreement.

Designing and Verifying Corrective Actions

Corrective action quality is where many RCAs collapse. The team identifies a plausible cause, then writes actions that are easy to assign but weak in effect. More training. Another checklist. A reminder email. A sign on the pump.

Those actions may have some value, but they rarely eliminate a recurring mechanical failure by themselves.

A professional engineering team conducts a collaborative root cause analysis meeting in a modern office workspace.

According to the SI Labs root cause analysis article, failure to implement systemic countermeasures over administrative ones results in recurrence rates exceeding 60% for pump and compressor failures, and structured fault trees with hierarchical layers are critical for eliminating breakdowns. That tracks with what reliability teams see in the field. The system usually has to change, not just the instruction.

Choose controls that change the system

For a pump seizure, stronger corrective actions usually sit higher on the safeguard hierarchy:

  • Design correction: Modify flush arrangement, bearing isolators, base stiffness, or minimum flow protection.
  • Control logic improvement: Restore trip logic, add permissives, or correct bypass valve control limits.
  • Component standardization: Replace marginal bearing or seal arrangements with a better-suited configuration.
  • Maintenance task redesign: Change lubrication method, alignment standard, or installation verification step.
  • Administrative support: Update procedure and training only after the physical and control issues are addressed.

One practical option for teams that need formal support is Forge Reliability's Root Cause Failure Analysis service, which applies structured methods such as Apollo, 5-Why, and fault tree logic to recurring equipment failures. In this context, it is one way to bring external rigor when internal teams are overloaded or too close to the event.

Verify before declaring victory

An action isn't complete because someone closed a work order. It is complete when the plant confirms that the failure mechanism has been removed or reduced to an acceptable level.

A verification plan should include:

  1. Success measure: What result proves the fix worked. Stable bearing temperature, reduced vibration pattern, restored hydraulic stability, or absence of repeated trips.
  2. Observation period: Long enough to expose the machine to normal operating variation.
  3. Responsible owner: Someone accountable for checking the result, not just implementing the task.
  4. Fallback action: What happens if the pilot or early deployment doesn't change the failure signature.

The corrective action should target the cause path that was validated, not the department that is easiest to assign work to.

For example, if a wastewater lift pump fails because ragging and low-flow operation combine to overload the impeller, then retraining operators to “watch the pump more closely” won't solve the issue. Installing or restoring minimum flow control, improving screening performance upstream, and modifying inspection intervals around known upset periods are stronger actions because they address the mechanism.

Good corrective actions also need budget and authority attached. A recommendation without funding is just a meeting note.

Documenting, Reporting, and Next Steps

RCA work only improves plant reliability when the findings become part of the operating system. That means the CMMS has to hold more than a closed work order and a vague failure code.

The investigation record should be easy to find, easy to audit, and useful to the next planner, engineer, or supervisor who inherits the asset.

Build the RCA record inside the CMMS

A practical CMMS-linked RCA package includes these elements:

  • Failure event record: Asset tag, date, operating mode, impact statement, and clear problem definition.
  • Timeline attachment: Alarm sequence, operator observations, work order history, and inspection timestamps.
  • Evidence file set: Photos, teardown notes, vibration snapshots, oil findings, process trends, and interview summaries.
  • Causal structure: The tree or map that shows how the team validated the cause path.
  • Corrective action tracker: Owner, due date, status, verification requirement, and closure evidence.
  • Failure code governance: Updated coding so the same event type can be trended later.

Plants that haven't tightened this workflow often discover that their CMMS stores symptoms, not learning. A better asset history starts with disciplined CMMS asset management and failure coding that supports analysis instead of hiding it.

Recurring pump seizure example

Consider a transfer pump in a chemical batching area that has seized multiple times over a year. Each event generated a repair work order. The descriptions read “bearing failed,” “pump locked,” and “motor overload.” None of those entries explained why the same unit kept failing.

A structured RCA would document the top event consistently, reconstruct the timeline across shifts, and connect the repair history to operating conditions. The team might find that the failures cluster after tank changeovers, when the pump runs at unstable low flow before the downstream valve line-up is completed. Teardown evidence could show repeated heat damage and contact marks consistent with poor hydraulic conditions, while inspection notes reveal that the minimum flow path wasn't functioning as intended.

The corrective plan would then be documented as a system change. Repair or redesign the minimum flow protection. Verify valve response. Add a startup permissive tied to line-up confirmation. Update PM tasks to inspect the bypass path. Assign each task to an owner with budget authority and track the verification period in the CMMS. That record is far more useful than another work order that says “replaced bearings.”

Turn one investigation into a reliability standard

The best RCAs don't stay isolated. They feed reliability standards across similar assets and sites.

A plant can turn one completed investigation into broader value by doing the following:

  • Standardize failure definitions: Make sure similar events use the same language and coding.
  • Replicate proven inspections: If seal flush verification matters on one pump train, apply that standard to the fleet.
  • Review PM content: Remove tasks that don't detect the validated failure mode and strengthen the ones that do.
  • Train on evidence quality: Teach supervisors and technicians what observations should be captured immediately after failure.
  • Audit action closure: Check whether corrective actions remain effective after turnover, staffing changes, or shutdown cycles.

Good reporting doesn't overwhelm leadership with detail. It gives leaders a clean summary of the event, the validated causes, the approved actions, and the current status of risk reduction. Reliability engineers need the technical depth. Plant leaders need confidence that the fix is real, funded, and being tracked.


Forge Reliability helps manufacturers and process plants turn recurring failures into documented corrective action plans through predictive maintenance, root cause failure analysis, and asset management support. Teams that want a clearer path to fewer repeat breakdowns can schedule a free reliability assessment with Forge Reliability.

Share this article

Rob Calloway

Rob Calloway

Rob Calloway is a Reliability Engineer and Condition Monitoring Specialist at Forge Reliability with 15+ years of experience in vibration analysis, root cause failure analysis, and integrated condition monitoring program development. He has worked across food & beverage, chemical processing, and manufacturing, helping maintenance teams catch developing equipment faults before they become unplanned shutdowns.

Get Started

Request a Free Reliability Assessment

Tell us about your equipment and facility. Our reliability team will review your situation and recommend a tailored reliability program — no obligation.

Free initial assessment
Response within 1 business day
No obligation or commitment

No obligation. Typical response within 24 hours.

Ready to Improve Your Plant Reliability?

Tell us about your facility and a reliability specialist will review your situation.

Claim Your Free Assessment →