By The SimplicityHub Team
A Failure Mode and Effects Analysis is a structured way to ask: what could go wrong, how bad would it be, and what are we doing about it? The output is a ranked list of risks, scored so teams work on the most dangerous ones first.
That sounds simple. In practice, most FMEAs fail for a specific reason: teams fill in the spreadsheet once during a project's Analyse or Improve phase, file it, and never look at it again. A good FMEA is a living document. It changes when processes change, when new data arrives, and when control measures prove themselves (or don't).
The three scoring dimensions
Every failure mode in an FMEA gets three scores, each on a 1 to 10 scale.
Severity measures the impact on the customer or downstream process if the failure occurs. A score of 1 means the effect is barely noticeable. A 10 means safety is compromised or regulatory requirements are violated. Severity is the one score that rarely changes, because the consequence of a failure mode is usually fixed.
Occurrence measures how frequently the failure mode is expected to happen. A 1 means it has essentially never been observed. A 10 means it is almost certain to occur in every batch or cycle. This score should come from data: defect rates, process capability indices (Cpk), or historical records. Guessing occurrence is one of the most common FMEA mistakes.
Detection measures the likelihood that existing controls will catch the failure before it reaches the customer. A 1 means current controls will almost certainly detect it. A 10 means there is no detection method in place. This is the dimension teams most often score incorrectly, because they confuse "we inspect for it" with "we reliably catch it."
Calculating and using RPN
Multiply the three scores together to get the Risk Priority Number (RPN):
RPN = Severity × Occurrence × Detection
The maximum possible RPN is 1,000 (10 × 10 × 10). The minimum is 1. Higher numbers get attention first.
A common threshold is to investigate any failure mode with an RPN above 100 to 125, but rigid cutoffs can mislead. A failure mode scored 10 (severity) × 2 (occurrence) × 5 (detection) produces an RPN of 100, while one scored 5 × 5 × 5 gives 125. The first one is more dangerous despite its lower RPN, because its severity is a 10. Always flag high-severity items regardless of RPN.
Some organisations now use the Action Priority (AP) method from the AIAG/VDA FMEA Handbook, which replaces the single RPN number with a three-level priority (High, Medium, Low) based on lookup tables for severity, occurrence, and detection combinations. The AP method addresses the multiplication problem described above.
Building the FMEA step by step
1. Define scope. An FMEA for an entire production line is too broad. Pick a specific process, subprocess, or component. A SIPOC diagram helps set boundaries.
2. List failure modes. For each process step, identify every way it could fail. Be specific. "Machine breaks down" is too vague. "Torque wrench fails to reach specified 25 Nm due to air pressure drop below 6 bar" gives the team something to act on.
3. Identify effects and causes. Each failure mode needs at least one effect (what the customer or next process step experiences) and at least one root cause. Use a fishbone diagram or 5 Whys to push past surface-level causes.
4. Document current controls. These fall into two categories: prevention controls (things that stop the cause from occurring) and detection controls (things that catch the failure after it occurs). List both.
5. Score and calculate RPN. Use your team's combined knowledge, but anchor scores in data wherever possible. If you have Cpk data for a process step, use it to inform the occurrence score. If your inspection records show a 2% escape rate, that detection score is not a 1.
6. Assign actions for high-priority items. Each action needs an owner and a target date. "Improve inspection" is not an action. "Install vision system on Station 3 by 15 September, validated against 500-piece sample" is.
7. Re-score after implementation. Once the action is complete, re-rate occurrence and detection. Severity stays the same (the potential consequence hasn't changed, only the likelihood). The revised RPN shows whether the action made a measurable difference.
Keeping the FMEA useful
Review the FMEA when any of these happen: a process change, a new supplier, a customer complaint related to a listed failure mode, or at a scheduled interval (quarterly is common in manufacturing). If the FMEA sits untouched between project phases, it is documentation, not risk management.
For teams working through DMAIC projects, the FMEA typically starts in the Analyse phase and carries forward into Control. The SigmaXL Excel add-in includes a DMAIC FMEA template with automated RPN sorting that handles the calculation and re-ranking as you update scores, which removes one friction point from keeping the document current.
Track the total RPN across all failure modes over time. A declining total RPN is one of the clearest indicators that a process improvement effort is reducing real risk, not just ticking boxes.
Common mistakes to avoid
Scoring by committee without data. Get the defect logs, Cpk values, and inspection escape rates into the room before you score.
Treating all RPNs equally. A failure mode with severity 9 and RPN 90 matters more than one with severity 3 and RPN 135. Look at severity first, RPN second.
One-and-done. If the FMEA only exists as a deliverable for a tollgate review, it is not doing its job. Build the review cycle into the control plan.
Vague actions. "Train operators" is not a corrective action. Specify who will be trained, on what, by when, and how you will verify it worked.