Reliance on AI drifts and nobody notices
Start here
Read the Real Outputs
Count what actually goes wrong
Read a large batch of real AI outputs as a team, write down what went wrong in each one, and count which problems happen most often.
Try this first
Pull 30 recent AI outputs. Have two people read them and write one line about each one that has a problem. Then count which problems repeat.
Borrowed from Hamel Husain, an engineer who advises companies on improving AI products, in his guide A Field Guide to Rapidly Improving AI Products. Read the original source.
Other moves for this problem (4 of 4)
Kind of change
- Check the Felt Speedup · Measure the time AI saves instead of estimating itFor when: Decisions about AI tools, staffing, or deadlines rest on how much time people say AI is saving, and nobody has measured it.
- Game days / chaos engineering · Break your own safeguards on purposeFor when: You don't know whether the team's safeguards would catch an AI failure.
- Track the Overrides · Keep a record of when people overrule the AI and who was rightFor when: People review AI recommendations before acting on them, but nobody knows how often they overrule the AI or whether their overrides turn out to be right.
- The Known-Answer SetFor when: Nobody can say whether the AI, or the team's trust in it, has gotten worse over the past few months.