Confirmation fatigue makes human approval a weak safety guard

Asking humans to approve every AI action does not produce safe behavior, because approval fatigue makes people rubber-stamp even dangerous prompts.


When people set up an AI assistant to take actions on their behalf — sending messages, editing files, making purchases, running commands — the most common safety instinct is the same: make it ask before it does anything. Simon Willison, a developer and writer who has spent years thinking about how AI tools fail, has a blunt assessment of that instinct:

Confirmation fatigue is real, and asking humans to click "OK" every few steps is clearly not going to result in safe behavior.

The argument is simple enough that most people have already lived a version of it. When a system interrupts you constantly for approval, you stop reading the prompts. The first few confirmations get real attention. The twentieth gets a reflexive click. The approval dialog becomes background noise — something between you and getting the thing done, rather than a moment where a judgment actually happens. And once approval becomes a reflex, it stops functioning as a guardrail at all. A dangerous request buried in a stream of harmless ones will sail through precisely because the harmless ones trained you not to look.

This is why the "always ask me first" setting — which feels like the cautious choice — may be less safe than it appears. It does not remove risk; it relocates it, onto a human attention span that the system's own behavior is steadily eroding. The more an assistant does, the more prompts it generates, and the faster your scrutiny decays.

Who this is for: anyone whose main safety control for an AI tool is their own approval click. That covers a lot of ground. Consumer assistants increasingly act on your behalf — booking, buying, replying — and many workplace tools gate risky actions behind a human sign-off. If that sign-off is your safety plan, Willison's point is that you should treat it as weaker than it looks. It is worth saying plainly that the observation originates in the developer world, where AI coding agents ask permission to run commands every few seconds and the fatigue sets in fast. But the mechanism is not developer-specific. Any high-frequency approval loop degrades the same way.

One honest caveat: this is an idea, not a tested prescription. Willison is describing a failure mode, not citing a study measuring how quickly people start rubber-stamping, and he is not offering a replacement design here. There is no announced product or feature that solves confirmation fatigue; it is an open problem in how these tools are built. So the practical takeaway is defensive rather than constructive: do not mistake an approval prompt for genuine oversight. If your assistant asks you to confirm things often, the risk is not just that you will approve something bad — it is that you will approve it without noticing it was bad, because the interface taught you that approving is what you do.

What does help, without being a complete fix, is reducing how often the question gets asked in the first place: letting an assistant act freely on low-stakes tasks and reserving your attention for the ones that are genuinely hard to undo — money moving, messages leaving, files being deleted. An approval you only see occasionally is one you might actually read. A constant stream of approvals is a guardrail that exists mostly on paper.

productsautomation