A maintenance manager watches the same centrifugal pump trip for the third time in eight weeks. One work order records a bearing replacement. Another records coupling realignment. The latest calls for a seal flush. Production returns each time, but the pump keeps failing under the same duty. The repair history looks active, yet the reliability problem remains untouched.
That pattern is familiar across pump, motor, and gearbox fleets. A failed component is visible and easy to replace, while the operating conditions, alarm response, lubrication practice, and maintenance strategy behind the failure are harder to examine. Apollo Root Cause Analysis gives the investigation a disciplined way to follow those interacting causes, test them against evidence, and connect the findings to condition monitoring.
Table of Contents
- Why Recurring Equipment Failures Force a Different Investigation Method
- What Apollo Root Cause Analysis Is
- Worked Example of a Centrifugal Pump Seizure Investigation
- Comparing Apollo RCA with 5-Why and Fault Tree Analysis
- Common Pitfalls That Undermine Apollo Investigations
- Connecting Apollo RCA to a Predictive Maintenance Program
- Templates, Checklist, and Next Steps for Your Team
Why Recurring Equipment Failures Force a Different Investigation Method
The third failure changes the question. The issue is no longer whether the bearing, coupling, or seal was damaged. The issue is why different repairs keep producing the same outcome. Replacing the last failed part can restore service, but it doesn't explain the pattern.
Single-cause thinking creates a narrow work order trail. A bearing failed, so the bearing was replaced. A coupling was misaligned, so the coupling was realigned. A seal leaked, so the seal was flushed. Each statement may be factually correct, but none proves that the listed component caused the full event.
The repeating failure itself is evidence that the investigation stopped too early. Complex events commonly involve multiple causes, and Apollo methodology was developed to challenge the assumption that one “root cause” explains everything. Its origins trace back to nuclear-industry problem solving after the 1979 Three Mile Island partial meltdown, when Dean L. Gano was working in the nuclear power sector and found conventional methods inconsistent. Training and government guidance describe the approach as developing in the mid-1980s and gaining broader visibility after Gano's 1999 book through this background on Apollo Root Cause Analysis.

The repair is not the investigation
A pump bearing can fail because of lubrication starvation, contamination, excessive load, shaft deflection, or a combination of conditions. The visible damage identifies the failure mechanism, meaning the physical process that produced the damage. It doesn't automatically identify the conditions that allowed that mechanism to develop.
A useful investigation separates three layers:
- Physical causes: Damage, loading, heat, wear, contamination, or component interaction.
- Human causes: Operating decisions, inspection quality, alarm response, or procedural execution.
- Latent system causes: Design assumptions, alarm settings, maintenance intervals, data quality, or resource decisions.
The practical distinction is central to defining contributing factors. A team that records only “bearing failure” has documented a symptom of the event. A team that shows how bearing damage combined with an undetected operating condition and a weak response process has built a prevention strategy.
Practical rule: A closed work order proves that someone performed a repair. It doesn't prove that recurrence risk was removed.
Apollo Root Cause Analysis is suited to this problem because it traces effects backward through a cause-and-effect structure. Each proposed cause must be necessary to the event and supported by evidence. If the chart cannot survive a challenge from operations, maintenance, or engineering, the action plan isn't ready.
What Apollo Root Cause Analysis Is
A centrifugal pump can return to service after a bearing change, only to fail again when the same operating conditions remain. Apollo Root Cause Analysis gives the investigation a way to explain that recurrence. It works backward from the undesired effect, mapping the conditions and relationships that produced it. Its RealityCharting logic tree allows several causal paths, instead of forcing the team into one linear answer, as described in this Apollo RCA introduction.
The method is associated with Dean L. Gano's Apollo approach. Training descriptions also identify Dean Bliss and Robert J. Latino among the experts who developed and advanced its practical use. For a plant team, the operating discipline matters more than the label. A cause enters the chart because evidence and causal logic support it, not because a meeting reached agreement.
Four stages that control the investigation
Define the problem. State what happened, where and when it happened, and the equipment or process state. “Pump failed” leaves too much uncertainty. “Cooling-water pump stopped after rising vibration, followed by higher motor current and shaft lockup” establishes a usable boundary.
Assemble the evidence pack. Preserve physical evidence and collect logs, inspection findings, maintenance history, operator accounts, condition-monitoring data, process records, and design information. One investigation workflow recommends securing the scene and assigning a multidisciplinary team within 24 hours, with the timing documented in the practical investigation workflow.
Construct the RealityChart. Begin with the effect, then add causes that explain it. Branches may include a physical failure, an operating condition, an inspection gap, or a system weakness. Show how causes combine. A list of possible faults is not a causal chart.
Verify causes and assign action. Test each link against facts and remove speculative branches. Assign actions where they can interrupt the chain, with an owner, due date, and effectiveness check. A four-stage summary describes the work as defining the problem, determining causal relationships, identifying effective solutions, and implementing and tracking them, as outlined in Apollo RCA training guidance.
Evidence gates separate Apollo from brainstorming
A pressure reading deserves scrutiny before it supports a conclusion. Check the instrument, calibration status, impulse line, and process context. Apply the same discipline to causal statements. If a team suspects misalignment caused a bearing failure, it should compare alignment records, shaft condition, vibration signatures, and damage patterns before keeping that branch.
Apollo investigations can take longer than a short meeting and corrective-action form. The trade-off is a defensible explanation that operations, maintenance, and engineering can challenge and trace. The chart should also support root cause failure analysis practices by connecting evidence to actions that can be checked after the repair.
The method rejects single-cause thinking. Government root-cause guidance emphasizes iterative tracing across connected causes and effects rather than stopping at the first plausible explanation, as noted in this Apollo RCA guidance document. For rotating equipment, that logic sets up the next step: verify the corrective action with vibration and oil analysis, not just a closed work order.
Worked Example of a Centrifugal Pump Seizure Investigation
Consider a centrifugal cooling-water pump that seized during normal production demand. The case is presented as a plant-floor example, not as a claim about a named facility. The investigation starts with the event, not with the damaged bearing.
Define the event before naming a cause
The event statement should capture the failure signature and sequence:
- The pump stopped delivering cooling water.
- Vibration increased before the trip.
- Motor current rose as the rotating assembly loaded.
- The shaft later seized.
- The timeline links the alarm, operator response, trip, and inspection.
That sequence matters. A seized pump isn't the same investigation as a pump with a slow seal leak. The event definition sets the boundary for evidence collection and prevents the team from mixing unrelated defects into the chart.
The next step is the evidence pack. For this pump, it would include:
- Oil analysis results and the wear-debris trend.
- The latest coupling alignment record.
- Vibration spectra and route history.
- Operator logs and alarm acknowledgements.
- The pump curve at failure compared with design expectations.
- Bearing, shaft, and seal inspection findings.
- Maintenance and work-order history.

Build branches that explain the seizure
The first RealityChart branch may be physical. Bearing damage could combine with shaft deflection, increasing friction and load until the pump seized. The team shouldn't accept that branch because the bearing is damaged. It needs matching evidence, such as wear patterns, vibration behavior, oil debris, shaft measurements, or a documented load condition.
A second branch may involve human action. The vibration alarm was acknowledged late, allowing the developing defect to continue. Interviews and alarm records can test that statement. The investigation should distinguish an operator's action from the system conditions that shaped the response.
A third branch may be latent. The alarm threshold was set too high for the failure mode, and the preventive-maintenance strategy ignored the oil-analysis trend. Those causes are verified through the alarm configuration, historical data, maintenance templates, and decisions governing route review.
Weak branches are pruned. A suspected process upset without supporting pump-curve or operating data stays outside the validated chart. A confirmed alignment deviation may remain relevant, but only if it is necessary to the seizure and linked to the observed damage.
Turn verified causes into prevention
Corrective actions should interrupt more than one point in the chain. The plant might correct the alarm strategy, revise the oil-analysis review process, inspect shaft condition, improve alignment controls, and define an escalation path for abnormal vibration. Each action needs a named owner and an effectiveness test.
The final test is not whether the chart looks complete. It is whether the pump's condition remains stable after the intervention. Vibration and oil analysis provide the verification route, while the maintenance system records whether the actions were completed and whether the failure signature returns. A detailed centrifugal pump root cause analysis approach helps keep the physical failure connected to the reliability decision.
Comparing Apollo RCA with 5-Why and Fault Tree Analysis
Maintenance managers rarely choose an investigation method based on theory alone. They consider the failure's complexity, the evidence available, the urgency of restoring production, and the skill of the team leading the work.
| Criterion | Apollo RCA | 5-Why | Fault Tree Analysis |
|---|---|---|---|
| Depth of causation | Follows multiple interacting physical, human, and system causes | Follows one primary chain of questions | Maps combinations of events leading to a defined top event |
| Evidence requirements | Requires evidence and causal logic for each accepted branch | Often depends on the quality of each answer | Requires credible failure logic and, for quantitative work, suitable probability data |
| Time investment | Moderate to high, depending on event complexity and evidence quality | Low for direct, well-bounded problems | High when the system has many gates, dependencies, or failure paths |
| Analyst skill | Needs facilitation, equipment knowledge, and evidence discipline | Accessible to general problem-solving teams | Requires strong systems and reliability analysis capability |
| Best-fit failures | Recurring, multi-factor equipment and process failures | Simple, linear, human-error-dominated events | Safety, design, and probabilistic system reviews |
Where each method works
5-Why works well when the failure has a clear single chain. A motor stopped because a protective device opened, the device opened because of an overload, and the overload can be tied to a specific, verified condition. The method becomes weak when several independent contributors interact, because the format encourages the team to choose one path.
Fault Tree Analysis starts with a top event and works backward through logical combinations. It excels during design reviews and risk studies where the team can define system dependencies and quantitative inputs. Incident investigations often lack the complete probability information needed for a meaningful quantitative fault tree, even when the logic itself is useful. Teams evaluating that method can review examples of fault tree analysis before selecting it.
Apollo RCA occupies a practical middle position for plant events. It handles multiple causal paths without requiring every incident to become a probabilistic model. Its evidence gates also make the result easier to defend during shift handovers, management reviews, and corrective-action audits.
The choice should match the problem. A short 5-Why may be responsible for a contained, linear defect. A complex pump seizure deserves a method that can preserve several verified paths without assigning blame to the first person in the sequence.
Common Pitfalls That Undermine Apollo Investigations
Apollo investigations fail for predictable reasons. Most breakdowns occur before the chart is complete, when production pressure pushes the team toward a convenient answer.
The first error is stopping at the first plausible physical cause. “The bearing failed” may be true, but it doesn't explain lubrication condition, loading, alignment, alarm response, or maintenance strategy. The second is skipping evidence gates. A chart built from assumptions can look logical while remaining impossible to verify.

Weak physical verification creates weak actions
Vibration spectra, oil wear debris, and thermal data must be matched to the suspected failure mode. A general statement that “vibration was high” doesn't establish bearing spalling, imbalance, looseness, resonance, or misalignment. The evidence needs to fit the mechanism and the timeline.
Teams also lose effectiveness when they assign vague actions:
- Owner: “Maintenance” isn't a person accountable for completion.
- Due date: An open-ended action can disappear during the next shutdown cycle.
- Acceptance criteria: “Improve monitoring” doesn't state what condition will prove success.
- Escalation path: A route technician needs clear instructions for what happens after an alarm.
Treat the chart as a live reliability record
An Apollo chart shouldn't be filed after the repair. It should change the next inspection route, alarm review, lubrication decision, or operating control. If the investigation closes without a measured reduction in repeat events or a defensible condition trend, the plant has completed paperwork rather than reliability improvement.
A finished chart with no effectiveness evidence is an untested hypothesis.
Blame also distorts the result. If an operator responded late to an alarm, the team still needs to examine alarm design, staffing, training, escalation rules, and the clarity of the operating procedure. Human action belongs on the chart, but blame rarely belongs at the end of it.
Connecting Apollo RCA to a Predictive Maintenance Program
A completed Apollo investigation becomes valuable when it changes what the plant measures. The RealityChart identifies failure paths, and the predictive maintenance program assigns a condition signal to each path.
For the seized cooling-water pump, the mapping might look like this:
- Bearing race damage: Use vibration routes that examine bearing-related energy, envelope behavior, and shock response.
- Lubrication starvation or contamination: Review oil-analysis wear trends and ferrography, which examines wear particles to help characterize active damage.
- Alignment drift: Schedule laser alignment checks and track thermal growth where operating temperature changes affect shaft position.
- Procedural shortcuts: Add behavior-based observations that examine how technicians and operators respond to alarms, lubrication tasks, and abnormal conditions.
The route must state more than the measurement type. It needs the asset, failure mode, collection frequency, trigger condition, and escalation owner. A vibration alarm should lead to a defined diagnostic review, not an unstructured note in the work-order system.
Convert corrective actions into monitored controls
A corrective action register can become a PdM control plan:
- Cause branch: Record the verified cause and the supporting evidence.
- Control point: Identify the condition that should reveal recurrence early.
- Route task: Define who collects the data and how the result is recorded.
- Trigger: State what change requires review or intervention.
- Escalation: Name the person responsible for deciding the next action.
- Effectiveness check: Compare the post-fix condition with the pre-fix baseline.
This approach prevents generic routes from hiding specific failure modes. A pump doesn't need “monthly vibration” solely because vibration is available. It needs monitoring designed around the damage mechanism identified in the chart. Guidance on predictive maintenance for manufacturing supports this failure-mode-specific view of condition monitoring.
Effectiveness can be assessed through stable vibration behavior, cleaner oil results, fewer repeat work orders, longer intervals between failures, and documented avoidance of the original mechanism. The selected measure should connect directly to the cause. A closed alignment action shouldn't be verified only by a signed form. It should be checked against alignment data, thermal behavior, and subsequent vibration.
Templates, Checklist, and Next Steps for Your Team
A reliability lead should assemble the Apollo investigation kit before the next shutdown. The kit must support evidence capture under production pressure, assign ownership clearly, and connect the findings to condition monitoring. That connection prevents the investigation from ending when the work order closes.
The basic investigation kit
RealityChart worksheet: Place the event at the top or center. Provide branches for physical, human, and latent system causes. Each cause entry should include its evidence reference and the causal relationship it supports.
Evidence pack template: Organize physical findings, operating logs, condition-monitoring trends, maintenance records, design or process data, and interview notes. Record the asset reference, collection date, source, and relevance for every item.
Corrective action register: Link each action to a chart branch. Include the action, accountable owner, due date, required resources, completion evidence, and effectiveness check.
PdM verification sheet: Record the failure mode, selected measurement, baseline condition, trigger criteria, route owner, escalation path, and post-action review date. For a centrifugal pump, this may connect a validated bearing or lubrication cause to vibration and oil checks.
Pre-investigation checklist
Before the team begins, the reliability or maintenance lead should confirm:
- Team selection: Include operations, maintenance, reliability, engineering, and an equipment subject-matter specialist.
- Scope freeze: Define the event boundary and separate the incident from unrelated defects.
- Evidence preservation: Secure damaged parts, photographs, samples, alarm records, and controller data before the scene changes.
- Data sources: Identify vibration, oil, temperature, current, alignment, process, work-order, and interview records.
- Gate criteria: Require evidence and causal necessity before validating a branch.
- Action control: Assign an owner, due date, completion evidence, and effectiveness measure to every accepted cause.
- PdM route: Decide how the route will detect recurrence before closing the investigation.
A field-by-field format
Use these fields for the event record:
- Event: Exact equipment failure and operating consequence.
- Timeline: Alarm, response, trip, inspection, and repair sequence.
- Observed effect: What the equipment did, without assumptions.
- Cause branch: Physical, human, or latent system condition.
- Evidence: Data, inspection, record, or testimony supporting the branch.
- Logic test: Why the cause was necessary to produce the event.
- Action: Intervention that interrupts the causal path.
- Owner and due date: Named accountability with a completion target.
- Effectiveness check: Condition or performance evidence required after implementation.
- Lesson transfer: Other assets, routes, procedures, or design standards requiring review.
This format keeps the investigation usable across shifts and sites. It also gives the team a practical decision point: use a full Apollo investigation when a pump, motor, or gearbox failure recurs and isolated repair has not removed the mechanism.
Templates provide a strong starting point, but complex repeat failures often need specialized facilitation, evidence review, and condition-monitoring experience. Forge Reliability applies Apollo, 5-Why, and fault tree methods with vibration analysis, oil analysis, thermography, ultrasound, and motor current signature analysis. Teams can request a free reliability assessment to connect recurring work orders with RealityChart branches and PdM evidence.
Forge Reliability can assess recurring pump, motor, and gearbox failures, then connect validated causes with practical monitoring routes and corrective actions. A free reliability assessment helps identify where the next repeat failure is most likely to start.