A practical walkthrough for running a Process FMEA — from assembling the right team to recalculating RPN after corrective actions are implemented.

Something has gone wrong. You are not sure when it started, or exactly where. You just know the defect reached the customer, and now you are in damage-control mode.
That is the problem FMEA is designed to prevent.
Failure Mode and Effects Analysis (FMEA) is a structured method for identifying what could go wrong in a process or product before it actually does. You map out every potential failure, assess how serious it would be, how likely it is to occur, and how likely your current controls are to catch it. Then you prioritise. Then you act.
The key word is "before." FMEA is a proactive tool. If you are using it only after a failure has occurred, you are using it reactively — which is what root cause analysis is for. The two tools are complementary, not interchangeable.
FMEA was originally developed by the US military in the 1940s and later adopted by the aerospace and automotive industries. Today it is used across manufacturing, healthcare, financial services, and any other environment where process failure carries a real cost.
This guide walks you through how to run one, step by step.
Before you start, be clear about which type you are running. The methodology is the same. The focus is different.
| Type | Focus | Typical Use |
|---|---|---|
| Design FMEA (DFMEA) | Product or component design | New product development, engineering changes |
| Process FMEA (PFMEA) | Manufacturing or service process steps | Production lines, transactional processes, service delivery |
| System FMEA | Interactions between subsystems | Complex systems with multiple interdependencies |
For most Lean Six Sigma projects, you will be running a Process FMEA. That is what this guide covers.
If you want a deeper look at how FMEA fits into the broader improvement toolkit, the SimplicityHub FMEA analysis guide covers the scoring mechanics and common pitfalls in detail.
A well-run FMEA follows a consistent sequence. Skipping steps — particularly the action-tracking and re-scoring steps at the end — is what turns FMEA into a paperwork exercise rather than a genuine risk reduction tool.
Do not run an FMEA alone. It is a cross-functional exercise by design.
Bring together people who actually touch the process: engineers, operators, quality staff, supervisors, and anyone whose work is directly affected by a failure in that process. Aim for four to seven people. Larger groups slow down scoring without improving accuracy.
Operators are essential. They know failure modes that never appear in documentation. If your FMEA team consists only of managers and engineers, you will miss things.
Before you open the FMEA form, map the process. Use a process flowchart or SIPOC diagram to define the steps clearly. Every team member needs to be looking at the same process before scoring begins.
This step is frequently skipped. It should not be. Teams that score without a shared process map end up with inconsistent failure modes and arguments about scope.
For each process step, ask one question: in what way could this step fail to perform its intended function?
Write each failure mode at the right level of specificity. "Machine fails" is too vague to act on. "Torque wrench applies insufficient torque during bolt tightening" is actionable. One process step can, and usually does, have multiple failure modes.
Be exhaustive at this stage. You can consolidate and remove duplicates later. The risk of being too narrow here is far greater than the risk of being too broad.
For each failure mode, you need two things:
The effect: What happens downstream if this failure occurs? What does the customer (internal or external) experience?
The cause: What triggers the failure in the first place?
Causes should be described in terms of something you can control or correct. "Operator error" is not a cause. "Insufficient torque applied due to uncalibrated tool" is.
A fishbone diagram is a practical tool for brainstorming causes systematically across categories such as equipment, materials, methods, and people.
This is where the FMEA form comes alive. Each failure mode is scored on three dimensions, each on a scale of 1 to 10.
| Score | Severity (S) | Occurrence (O) | Detection (D) |
|---|---|---|---|
| 1-2 | No effect / minor | Remote likelihood | Almost certain to detect |
| 3-4 | Minor effect | Low likelihood | High likelihood of detection |
| 5-6 | Moderate effect | Moderate likelihood | Moderate likelihood of detection |
| 7-8 | Significant effect | High likelihood | Low likelihood of detection |
| 9-10 | Catastrophic / safety risk | Very high / near certain | Cannot detect |
Score as a team, not as individuals. When one person scores a 3 and another scores an 8, that disagreement matters. It usually reveals a different understanding of the process, the customer, or the controls in place. Work through those gaps before moving on.
Note on Detection scoring: A high Detection score (8, 9, or 10) means your current controls are unlikely to catch the failure. It does not mean the failure is easy to detect. The scale runs counter-intuitively: 1 is good, 10 is bad.
Multiply the three scores together:
RPN = Severity (S) × Occurrence (O) × Detection (D)
The maximum possible RPN is 1,000 (10 × 10 × 10). In practice, most teams set an action threshold somewhere between 100 and 125. Any failure mode above that threshold requires a corrective action.
However, do not rely on RPN alone. A failure mode with a Severity score of 10 deserves attention regardless of its RPN. If a failure is catastrophic or poses a safety risk, it is a priority. Full stop.
Sort your FMEA table by RPN descending. Focus your team's energy on the top of that list.
For each high-priority failure mode, assign:
A specific corrective action (not "investigate further")
A responsible owner
A target completion date
Corrective actions should aim to reduce Occurrence (fix the root cause so the failure happens less often) or improve Detection (add a control so you catch it earlier). Reducing Severity typically requires a design change, which is harder to achieve in a process FMEA but not impossible.
Connect this step to your improvement cycle. FMEA tells you what to fix. Your DMAIC or PDCA project is how you fix it.
Once actions are completed, re-score the affected failure modes. Severity rarely changes. Occurrence and Detection should improve if your actions worked.
Recalculate the RPN. Document the new scores alongside the original ones.
This final step is what separates a living FMEA from a one-time document. If you never recalculate, you have no evidence that the risk was actually reduced. You just have a form with a date on it.
Most FMEA failures are not failures of the method. They are failures of execution. These are the patterns to watch for.
An FMEA completed at project launch and never revisited is not a risk management tool. It is a historical document. Processes change. Equipment changes. People change. Your FMEA should be reviewed whenever a significant process change occurs, not just at initial launch.
"Equipment malfunction" is not a failure mode. "Seal degradation in Pump 3 causing hydraulic pressure loss below 40 bar" is. The more specific your failure modes, the more useful your actions will be. Vague entries produce vague actions, which produce no improvement.
Severity, Occurrence, and Detection scores are only as good as the knowledge of the people assigning them. A quality manager scoring Occurrence without input from the operator who runs the process daily will get it wrong. Bring the right people. Score together.
An RPN of 80 sounds manageable. But if that 80 comes from S=10, O=2, D=4, you have a catastrophic failure mode that occurs rarely but is almost certain to reach the customer when it does. That is not manageable. Treat Severity 9 and 10 scores as automatic priorities, regardless of RPN.
The action plan column of an FMEA form is where most FMEAs die. Actions are assigned, dates pass, and nobody recalculates the RPN. Build a review cadence into the project. Treat the FMEA like any other project action log: someone is accountable, there is a due date, and completion is verified.
FMEA is most commonly used in the Analyse and Improve phases of a DMAIC project.
In the Analyse phase, it helps you structure your thinking about potential failure causes before you commit to a root cause. In the Improve phase, it validates that your proposed solution does not introduce new failure modes while addressing the original problem.
It also has a role in the Control phase. A revised FMEA, updated to reflect the improved process, forms part of your control plan. It documents the residual risks and the controls in place to manage them.
FMEA does not replace root cause analysis. It is a risk prioritisation tool, not a diagnostic one. Use it to decide where to focus your improvement effort. Use root cause analysis tools — 5 Whys, fishbone diagrams, fault tree analysis — to dig into the causes once you have identified the highest-priority failure modes.
The two approaches are most powerful when used together: FMEA to prioritise, root cause analysis to diagnose.
You do not need specialist software to run an effective FMEA. A well-structured spreadsheet is sufficient for most process improvement projects.
What you do need is the right structure: process steps listed clearly, failure modes written at the right level of specificity, scores agreed as a team, and an action plan with named owners and dates.
Download the free Failure Mode Identification Template to get started. It includes the scoring framework, RPN calculation, and action tracking columns you need to run a complete FMEA from first session to close-out.
If you want to build the broader skills to lead FMEA within a structured improvement project, the SimplicityHub Black Belt programme covers FMEA in depth alongside Design of Experiments, hypothesis testing, and full DMAIC methodology.
The form is straightforward. The discipline is not. Run it properly and FMEA will tell you exactly where your process is most at risk — before your customers find out for you.
A Process FMEA is proactive — it identifies what could go wrong before it happens. Root cause analysis is reactive — it investigates a failure that has already occurred. The two are complementary: use FMEA to prioritise risk, then use root cause tools like 5 Whys or a fishbone diagram to diagnose issues once they surface.
Aim for four to seven people, including operators who actually run the process, not just engineers and managers. Operators surface failure modes that rarely appear in documentation. Larger groups slow down scoring without improving accuracy.
Most teams set an action threshold between 100 and 125 out of a possible 1,000. However, RPN should never be the only trigger — any failure mode with a Severity score of 9 or 10 deserves attention regardless of its overall RPN, because the consequence of failure is catastrophic.
A low RPN can hide a high-Severity risk. For example, Severity 10, Occurrence 2, Detection 4 gives an RPN of only 80, but a Severity of 10 means the failure is catastrophic when it does occur. Treat Severity 9 and 10 scores as automatic priorities, independent of the total RPN.
Re-score the affected failure modes and recalculate the RPN. Occurrence and Detection should improve if the action worked; Severity rarely changes. Without this recalculation step, you have no evidence the risk was actually reduced — just a form with a date on it.
A ready-to-use failure mode template with scoring, RPN calculation and action tracking built in.