Learn and improve
Read the Real Outputs
Count what actually goes wrong
Read a large batch of real AI outputs as a team, write down what went wrong in each one, and count which problems happen most often.
Borrowed from Hamel Husain, an engineer who advises companies on improving AI products, in his guide A Field Guide to Rapidly Improving AI Products. Read the original source.
One session · An hour or more to set up · Anyone on the team can start it · Start here for this problem
Try this first
Pull 30 recent AI outputs. Have two people read them and write one line about each one that has a problem. Then count which problems repeat.
Use it when
The team relies on AI output every week, but nobody has looked closely at a large sample of it in months.
Skip it when
Skip it when the team uses AI only occasionally, or when there is no way to gather a batch of real outputs. It also matters less for work that is already checked line by line before it goes anywhere.
How to introduce it
Pull 50 to 100 recent examples of AI output from real work. These can be drafts, summaries, answers to customers, or code. Split them among the team. For each example, write a short note in plain words about anything that is wrong or weak. Do not use a checklist yet. A checklist only finds the problems you already expect. When everyone is done, put the notes side by side and group similar ones together. Count how many examples fall into each group. Fix the largest group first. Then repeat the exercise a month later to see whether that group got smaller.
How to show up
Keep the notes descriptive. Ask people to write what they see, such as "the date is wrong" or "it ignored the customer's question." Discourage guesses about why the AI did it. Hold off on fixes until the counting is done, because the first problem people notice is often not the most common one.
How long it takes
About two hours for the first round with 50 to 100 examples, including sorting and counting. Later rounds are faster, because the problem types already exist.
What makes it hard
The reading is slow, and people want to start fixing things after the first few examples. Gathering the outputs can also be harder than expected when people use AI through personal accounts. Decide in advance where the examples will come from.
What it looks like when it's working
The team can name its top two or three AI problems and roughly how often each one happens. After a fix, the count for that problem goes down in the next round.
How long until it sticks
Two or three rounds, about a month apart, before the team starts gathering examples without being asked.
How you know it stuck
The team pulls a fresh sample on a regular schedule, and new problem types get added to the list as they appear.
The idea behind it
Teams usually form their opinion of an AI tool from a handful of memorable examples, either very good or very bad. Those examples are not a reliable guide to what happens most of the time. Reading a large sample and counting the problems gives the team an accurate picture of what goes wrong and how often.
Where it comes from
AI evaluation practice. This comes from the way engineers who build AI products find out what is actually going wrong with them. Hamel Husain has advised more than 30 companies on this. He found that teams usually guess wrong about their AI's main problems until they read real examples. In one case, three types of problem caused more than 60 percent of the failures. Fixing them raised success on one task from 33 percent to 95 percent. The same idea has a long history in quality management. Most defects usually come from a small number of causes, and counting them shows which causes matter. The team reads a large batch of real AI outputs together. Each person writes down what went wrong in each one. The team then sorts the problems into types, counts how often each type appears, and fixes the most common type first.