Learn and improve
The Failure Budget
Decide in advance how much failure is acceptable in AI experiments, so the team can take real risks instead of freezing.
Borrowed from error budgets in site reliability engineering, as described in Google's Site Reliability Engineering book. Read the original source.
Changes a rule or role · Minutes to start · Needs senior backing
Try this first
Agree one number before a round of experiments. How much failure is acceptable here?
Use it when
The team freezes at the first error and has no room to experiment.
Skip it when
This is for experimentation. Safety-critical and irreversible work gets no failure budget. Match it to contexts where trying and failing is the point.
How to introduce it
Perfect reliability is the wrong target for experiments. Decide in advance how much failure you will accept and treat it as a budget the team is allowed to spend on risk. Permission to fail, with a number on it. The team can then try bold things instead of freezing the first time something breaks.
How to show up
Set an honest rate, then honor it. Failures inside the budget are expected costs, not problems. Let one draw blame and the fear comes straight back.
How long it takes
A short conversation to set the budget before a round of experiments.
What makes it hard
Naming an acceptable failure rate feels like endorsing failure, so leaders resist it. Frame it as the price of learning. And the budget is only real if failures inside it pass without blame.
What it looks like when it's working
The team takes real risks and treats a failure inside the budget calmly. Playing safe, and alarm at every failure, means the budget is words.
How long until it sticks
Expect a few cycles before the team trusts that within-budget failures are acceptable.
How you know it stuck
The team experiments inside an agreed budget and a failure inside it produces learning.
The idea behind it
Zero failures means zero real risks, which means nothing learned. A budget turns failure from a thing to fear into a resource to spend.
Where it comes from
Blameless postmortem practice in engineering. Pre-allocate an acceptable failure rate for AI experiments so the team can take real risks instead of freezing at the first mistake. Since 100% reliability is the wrong target, the gap below it is a budget the team may spend on risk and experimentation, permission to fail, quantified.
Why it matters for senior teams
Only a senior team can install it
Also fits: Everything escalates to the top