AI that proves it did what it claims

This release's entire theme is that the system now proves what it claims, attacking the oldest AI failure mode of saying "done" when it isn't.


The newest release of Daniel Miessler's AI assistant setup is built around a single idea: the system should not just say it did something — it should be able to show it. In Miessler's words:

the system now proves what it claims

That may sound like a small distinction. It is not. The oldest failure mode in AI assistants is confidently reporting success on work that was never done, was half-done, or was done wrong. Anyone who has used these tools for more than a week has hit it: you ask for something, the assistant says done, and when you check, the file was never created, the email was never sent, the test never ran. The assistant is not lying exactly — it is predicting that the task probably went fine, and stating that prediction as fact.

This release attacks that failure directly by making verification part of the system itself rather than something you, the human, have to supply by double-checking everything. Instead of trusting the assistant's summary of its own work, the design is that each claim comes with evidence — the actual output, the actual result, something you can inspect.

Who this is for. Honestly, this is material for people who run AI assistants on real, consequential work — and a meaningful slice of it is aimed at people technical enough to build or configure such a system, which today still skews toward developers and serious hobbyists. If you use an assistant casually for drafting and brainstorming, the failure mode this fixes is real but less dangerous for you; a wrong paragraph is visible on its face. Where this matters most is when the assistant acts on your behalf — sending things, changing things, completing multi-step tasks — and its word is the only thing standing between you and a silent mistake. That is the situation where "trust but verify" collapses into just "trust," because verifying everything yourself defeats the point of delegating.

Why it matters. As assistants get handed longer and more autonomous tasks, the gap between "said it was done" and "was done" becomes the main risk. A system that produces proof alongside its claims changes your job from re-doing the work to auditing it — a much smaller job.

Is it usable today? Yes — it is shipping, not a proposal or a talk about the future. Two honest caveats, though. First, what "proof" looks like in practice, and how much setup it takes, is not spelled out in the announcement — the claim is that the system verifies itself, not that verification is effortless or universal across every kind of task. Second, proof of execution is not proof of correctness: a system can genuinely show it ran the steps it claims and still have done the wrong thing. This closes the lying-about-done gap, which is the biggest one, but it does not remove the need for judgment about whether the work was the right work.

If you are delegating real tasks to an assistant, the standard this sets — evidence over assurance — is the right one to hold any system to, whether or not you use Miessler's.

securityaccuracyautomationdeveloper
Source: github.com