Claims now declare what kind of evidence closes them, and the machinery routes each claim to a verifier that can actually produce that evidence, giving the ban on "should work" enforcement machinery.
"Should work" is the two-word epitaph of a lot of bad work done by AI assistants. Ask whether the file was actually saved, whether the tests passed, whether the email went out, and the answer is often a confident paraphrase of the request rather than a check of reality. Daniel Miessler has described a mechanism designed to put that habit under enforcement:
Claims now declare what kind of evidence closes them, and the machinery routes each claim to a verifier that can actually produce that evidence.
In plain language, this works like a sign-off rule at a workplace. Normally, an assistant makes a statement — the report is ready, the bug is fixed, the data was cleaned — and nothing in the system defines what would prove it. Under this approach, every claim carries a label saying what counts as proof: a passing test run, a file that exists with the right contents, a confirmation from the system that sent the message. Then the claim is automatically handed to whatever tool or check can actually produce that proof, rather than to the assistant itself, which would otherwise just assert again with more confidence. The claim does not close until the evidence arrives.
This matters because the core weakness of current assistants is not that they lie — it is that they conflate intention with outcome. They describe what was supposed to happen as though it did happen, in fluent, unhedged language. A system that forces each claim to name its proof, and then sends the claim to something that can verify it, turns the ban on "should work" from a polite instruction in a prompt into actual machinery. It is the difference between asking a contractor to promise the plumbing works and requiring a photo of water running.
Who is this for? Anyone who delegates tasks to an AI and currently spends effort re-checking its claims — which is to say, almost every serious user. But the honest caveat is that the implementation Miessler describes is engineering. Building claims that declare their evidence type and routing them to verifiers is the kind of thing done inside agent frameworks and orchestration code, largely developer territory. A non-developer is unlikely to switch this on themselves today; what they can take from it is the principle. If you find yourself accepting assurances, you can already mimic the idea in a cruder way: instead of asking did it work, ask for the specific artifact — show me the output, paste the confirmation, list the files it created. Requiring a named proof rather than a summary is the same discipline, applied manually.
On availability: Miessler describes this as shipping — it exists as working machinery, not a proposal. What is not stated is where it ships or how an ordinary user would reach it. There is no product name, pricing, or setup path in what he has said, so treat it as a capability that exists in his tooling rather than a feature you can assume is in whatever assistant you already use. It also does not make verification foolproof: a claim can be routed to a verifier and still be verified against the wrong thing, if the evidence type was declared loosely. Declaring what counts as proof is itself a judgment call, and a badly chosen one just gives you confident wrongness with paperwork attached.
The deeper idea, though, is portable and worth holding onto: an assistant's assertion should be treated as a hypothesis, not a result. Whether the enforcement is automatic or something you demand in your prompts, the standard is the same — a claim is only done when the evidence says so.