Living FMEA with AI — the risk map that updates with what actually happens on the line, not with what was estimated two years ago.

The FMEA is built when the process launches and dies in an Excel file: the real line events — a new defect, a drift, an incident dictated by voice — never make it back to the analysis. iLEAN links every captured event to its failure mode, recalculates the observed occurrence against the estimated one and proposes the RPN update with the evidence behind it. The person validates and signs. The FMEA stops being an audit document and becomes the plant's living risk map.

‹ See all Lean methods

Quality manager reviewing the plant's risk dossier with the living FMEA updated from real line events for an IATF audit
The problem

The FMEA closes on launch day. The line keeps talking for years, and nobody answers it.

Failure Mode and Effects Analysis is one of the best tools industrial engineering has produced: it forces a cross-functional team to sit down before producing and imagine every way the process can fail. The problem isn't the methodology. The problem is what happens after the launch:

  1. It closes as a deliverable — the FMEA is built to pass the industrialization milestone. The day production starts, the document is archived and the team disbands.
  2. Occurrence is an estimate that is never tested — the number feeding the RPN came out of the team's judgment in a meeting room. Two years later, the plant has thousands of real events that would say otherwise, and nobody has cross-referenced them with the document.
  3. Real events have no path back — the defect the operator logs, the drift Edge detects, the incident someone dictates by voice at the end of the shift: all of it ends up in a system separate from the FMEA's. Two worlds that don't talk to each other.
  4. It updates only when someone outside forces it — a product change, a customer claim or the audit notice. And then it gets updated in a rush, from memory, so the dossier adds up.

The result is a document that says one thing and a plant that does another. The living FMEA doesn't change the methodology: it gives it back the return channel it never had.

How it fits the IRIS system

iLEAN doesn't replace the FMEA — it wires the line into it, live.

For an FMEA to be alive you need two things that until now were expensive: capturing all plant events with enough structure, and having someone who reads and classifies them one by one without tiring. Edge and Connect solve the first; the agents solve the second. Connect is the putty that fills the cracks between the MES, the defect log, the voice-dictated shift report and the FMEA sheet.

The quality manager no longer arrives at the review defending estimates from two years ago. They arrive with the observed occurrence, the evidence that supports it and an RPN proposal on the table.

The iLEAN pieces applied to the living FMEA:

  • Edge — machine vision and in-line sensors: every detected defect is an event with an image, a reference, a shift and an exact moment. It works without a network.
  • Connect — collects what already exists and goes unused: MES rejection records, SCADA alerts, shift reports, incidents dictated by voice at the terminal or through the plant's messaging channel. Without this capture, the living FMEA remains an FMEA with better formatting.
  • Agents — the brain. They use the FMEA's structure as a target dictionary, link every event to its failure mode, detect the events that fit none (candidates for a new failure mode), calculate the observed occurrence and compare it with the estimate, and prepare the RPN update proposal with its evidence file.
  • Three safety rings — no figure in the document changes on its own. The agent proposes and attaches proof; the FMEA owner validates, corrects or discards, and their signature is recorded.

See the full IRIS architecture →

Before and after

Document FMEA vs. living FMEA with iLEAN

AspectClassic document FMEALiving FMEA with iLEAN
Review frequencyAt launch and when an audit or a change demands itContinuous: every captured event feeds the analysis
Source of the occurrenceTeam estimate during industrializationObserved on real events, tested against the estimate
New failure modesDiscovered when the customer claim arrivesThe agent proposes them on detecting events that fit none
RPN recalculationManual, in a meeting, from memoryProposed with evidence attached; the person validates and signs
Audit traceabilityReconstructed before the visitGenerates itself: event, date, validator, associated action
Link to the control planThey drift apart over timeEvery RPN change flags the affected control
Impact estimate

Impact estimate for your plant — to be validated with your numbers.

The block below is an estimate to be validated with the specific data of your plant. We put it forward so the committee has an order of magnitude; we refine it during the diagnostic.

  • Plant with a formal PFMEA per product family, reviewed reactively (product change, claim or audit) and with the occurrence never tested against the accumulated real events.
  • Pilot on one line where event capture is already running (Edge and/or Connect): event-to-failure-mode linking, observed occurrence and RPN proposal. First value expected within a few weeks.
  • The living FMEA is an enabling piece, not a direct-savings piece: its value is strategic. It enables the improvement cycle — because it turns every event into information that prioritizes where to act — and it underpins the IATF and APQP dossier with evidence generated in the daily work itself.
  • Expected reduction in the time spent preparing the FMEA review and the audit evidence on the order of ≥40% (conservative estimate), plus the indirect effect of catching earlier the failure modes the original analysis didn't foresee.

And the quality manager's reasonable doubt

“What if the AI misclassifies an event and puts noise into my document?” — hallucination is a problem of free generation, not of anchored tasks. Linking a described event to one of the failure modes that already exist in the FMEA itself is a textbook anchored task: the target catalog is closed, it's written down, and the model doesn't invent it. In this kind of task the best models are below 1.5% error [1]. And even so, the link arrives as a proposal with its confidence level, and nothing enters the document without the owner's signature. The three safety rings are there precisely for this.

[1] OpenAI paper “Why Language Models Hallucinate”, 2025 — on the reliability of AI in anchored tasks.

You may also be interested in: Real-time SPC with AI · AI-assisted measurement MSA · ISO 9001 with continuous audit and AI

Frequently asked questions

What people ask about the living FMEA with AI

Why does the FMEA die in an Excel file a few months after the process launches?

Because the FMEA is born as a deliverable, not as an operating tool. It's built during industrialization, with a cross-functional team meeting for weeks, and it closes the day production starts. From then on, the line generates real events every day — new defects, drifts, incidents dictated by voice, stoppages — but none of those events has a path back to the document. Updating it means convening the team again, and that only happens when an audit or a product change demands it. Result: the occurrence estimated two years ago still rules the RPN while the plant's reality is already something else.

How is a real line event linked to its FMEA failure mode?

The agent starts from the FMEA's own structure — process, operation, failure mode, effect, cause — and uses it as a target dictionary. When an event comes in (a defect detected by Edge, a rejection logged at the terminal, an incident the operator dictates by voice), it normalizes it and proposes which failure mode it corresponds to, with its confidence level. If the event fits none, it flags it as a candidate for a new failure mode — which is precisely the most valuable finding, because it means the process is failing down a path the original FMEA didn't foresee. The quality engineer confirms, corrects or discards the link.

Who validates the new RPN? Does the AI change it on its own?

No. The agent proposes; the FMEA owner validates and signs. What the system does is calculate the observed occurrence from the accumulated real events and set it beside the occurrence estimated back in the day, with the evidence behind it: how many events, on which dates, on which shifts, with which batch or reference. The RPN proposal arrives with that file attached, so the review meeting argues over data rather than memories. The change only enters the document when a person with authority approves it, and who, when and on what justification is recorded.

How does a living FMEA help in an IATF 16949 or VDA audit?

The auditor doesn't ask whether you have an FMEA — they ask whether it's a living document and whether the derived actions get closed. That's where the dossier usually breaks: there's an FMEA, there's a control plan, but there's no traceability between what happened in the plant and the revision of the analysis. With the living FMEA, every occurrence update carries the events that motivated it, the date, the validator and the associated containment or improvement action. The evidence that the loop is closed generates itself, at the very moment the work is done, instead of being rebuilt in a rush the week before the audit.

Does it work with design FMEA (DFMEA) and process FMEA (PFMEA)?

The natural, immediate-value case is the PFMEA, because the events captured in the plant are process events and the link is direct. With the DFMEA the loop is longer but it works too: when a process failure mode keeps repeating and the analysis concludes the root cause is a design cause, the system escalates it as feedback to the DFMEA and to the product engineering team. It's the loop APQP methodology always described and that in practice almost never gets walked, because nobody had the aggregated data to justify it.

Let's talk

Tell us your case and in 48h we'll send you the estimated ROI of putting your FMEA live.

We work on your plant's real data, not ours. Diagnostic with no commitment.

Request estimated ROI in 48h ‹ See all Lean methods