Learn and improve
The Known-Answer Set
Cases with known answers, held back and run periodically to catch drift in the model and in the team.
Developed by Superadditive in our own review of the research. It draws on machine-learning evaluation practice, where a held-out set of cases with known answers shows when a model gets worse.
Ongoing habit · An hour or more to set up · Anyone on the team can start it
Try this first
Set aside a handful of known-answer cases and run them now and then, against the model and against the team.
Use it when
Nobody can say whether the AI, or the team's trust in it, has gotten worse over the past few months.
Skip it when
You need a domain with knowable answers you can hold back. Where the work is subjective this does not apply, and a stale or tiny set gives false comfort.
How to introduce it
Keep a small set of cases where you know the right answer, held back from daily work. Run the model against them from time to time, and run the team against them too. It catches two failures that give no warning: the model degrading as conditions shift, and the team's own unaided skill going soft.
How to show up
Help build the set, then protect the schedule. The value is entirely in running it, so guard the cadence against everything that will crowd it out.
How long it takes
Some effort to build the set. Short, regular runs thereafter.
What makes it hard
Nothing looks wrong, so this is always the maintenance that gets postponed. Protecting the cadence is the challenge. Keep the set out of daily use and refresh it, because a leaked or stale set stops catching drift.
What it looks like when it's working
The runs catch drift or skill erosion before a real failure, and the team acts on what they show. A set that exists and never runs protects nothing.
How long until it sticks
Expect a few cycles before running against known answers becomes a protected routine.
How you know it stuck
The team runs the check even when nothing seems wrong.
The idea behind it
Both the model and the team can decay silently until a failure exposes it. A held-back set is your early warning.
Where it comes from
our review. The AI silently degrades when conditions shift, and the team's unaided skill quietly erodes, with no signal until something breaks. Keep a small held-out set of cases with known-good answers and periodically run both the AI and the team's reliance on it against them, to detect drift and skill atrophy before they bite.
Evidence
No outside source. This move came out of our own review of the research, so treat it as reasoned judgment rather than tested practice.
Also fits: Warning signs get explained away
Home problem: Reliance on AI drifts and nobody notices