Learn and improve
The 5 Whys, live
Run a root-cause walkthrough on the spot when AI output fails, instead of patching the symptom and moving on.
Borrowed from the Toyota Production System, where asking "why" five times is a standard way to find the root cause of a problem. Read the original source.
One session · Minutes to start · Anyone on the team can start it
Try this first
Ask why it happened. Then ask why of that answer. A few times.
Use it when
The same failure keeps recurring and teams keep patching symptoms.
Skip it when
Save it for failures worth understanding, not every glitch. And keep it blame-free, because the moment the whys start hunting for a culprit, honest answers stop.
How to introduce it
When something a model produced goes wrong, do not patch it and move on. On the spot, with whoever caught it, ask why it happened. Then ask why of that answer. A few rounds gets you under the symptom to the reason it was possible.
How to show up
Run it live, right when the failure happens, and keep it on the system rather than the person. Push past the first easy answer, because the root sits two or three rounds deeper than where people want to stop.
How long it takes
Ten to twenty minutes, live, when a failure occurs.
What makes it hard
People want the immediate fix and to move on, so the pause takes discipline. The chain slides toward blaming a person too, so keep steering it back to what in the process allowed it.
What it looks like when it's working
Failures produce fixes that prevent a recurrence rather than patches. The same failure returning means you are working the surface.
How long until it sticks
Expect a few uses before digging to the root becomes the reflex over quick patching.
How you know it stuck
People trace an AI failure to its root instead of patching it.
The idea behind it
The first explanation is almost always a symptom. A few rounds of why gets past the patch to the thing that let it happen.
Where it comes from
Toyota Production System. A root-cause ritual instead of a quick patch when AI output fails. When the cord is pulled, the team runs a root-cause walkthrough on the spot, with the person who spotted the problem.