A more robust way to run agents is to give them no access to data or tools that could cause harm if triggered wrongly.
Simon Willison, a developer and longtime writer about AI tools, recently described where his own thinking on agent safety is heading:
I'm personally inspired to double down on figuring out a productive way to run agents such that they don't have access to data or tools that can cause harm if triggered in the wrong way.
That is a statement of intent, not a product announcement. Nothing shipped. But the principle underneath it is worth understanding, because it applies whether or not you ever write a line of code.
An AI "agent" is software that can act on your behalf — read your email, browse sites, run commands, move files — rather than just answering questions. The safety problem with agents is not only that they make mistakes. It is that they can be manipulated. If an agent reads a malicious email or webpage, that content can contain instructions the agent might follow — the attack pattern known as prompt injection. Guardrails and better prompting reduce the odds, but nobody has eliminated them.
Willison's framing sidesteps that problem entirely: instead of trying to make the agent perfectly trustworthy, make the environment it operates in incapable of harm. If the agent has no access to your bank account, no permission to send email, no ability to delete files, then even a fully hijacked agent has nothing to grab. The lock matters more than the lockpick-resistance of the person holding the keys.
The brief says this points to a principle non-developers can apply, and that is mostly true — but with a caveat. Willison himself is talking about how developers run agents, and there is no ready-made product here. What a non-developer reader can take away is a question to ask of any AI tool that acts for you: what can this thing actually reach? Does the assistant connected to your calendar also have your inbox? Can the tool drafting replies send them without you? The answer is a settings question, not a coding question, and "give it less access" is a decision you can make today in whatever permissions screen the tool gives you.
The harder version — building genuinely isolated environments where an agent literally cannot touch anything harmful — is developer work, and it is honest to say so. That is the part Willison is still figuring out.
Treat this as an idea people are actively discussing, not a solved problem. Willison says he is inspired to figure out a productive way to do it — the word "productive" is doing real work there, because the tension is obvious: an agent that can reach nothing useful is useless, and an agent that can reach useful things can reach harmful ones. Where to draw that line, per tool and per task, is unresolved.
One thing a vendor pitch would not mention: this framing implicitly concedes that prompt injection is not going to be fixed by smarter models alone. If the answer were just "make the AI better at refusing bad instructions," you would not need to remove access in the first place. The design principle exists precisely because the failure mode is assumed.
If you use AI assistants that take actions, the practical takeaway is modest but real: audit what each one can touch, and prefer the narrowest access that still gets the job done.