AI agents can build a working software tool from a short brief

Simon Willison handed two AI coding agents a short research-spike spec and they produced a new open-source library good enough to release as an alpha with very few follow-up prompts.


This morning, Simon Willison — a well-known software developer and one of the closest watchers of what AI tools can actually do — handed two AI coding agents a short specification and let them build a piece of software. He described it as a "shower project": the kind of idea that occurs to you and that, until recently, would have stayed an idea unless you had hours of programming time to spend on it.

This morning (literally a shower project) I tasked Codex and GPT-5.6 Sol Ultra with building a prototype:

The result was a new open-source library — a small, reusable piece of code that other developers can drop into their own projects — released as an alpha, meaning an early public version that works but may still change. What is striking is not that an AI produced some code, which has been possible for a while. It is that the agents took a short written brief — a "research spike," roughly an outline for exploring whether an idea is feasible — and turned it into a working, tested, releasable tool with very little steering:

It took very few follow-up prompts to produce this project in a state good enough to release as an alpha.

A few things are worth unpacking for a non-developer reader. "Agents" here means AI tools that do more than chat: they can create files, run tests, notice failures, and fix them — working toward a goal over many steps rather than answering a single question. "Open source" means the result is public and free for anyone to inspect or use, so this is a verifiable artifact rather than a private demo.

Who is this for? Honestly, the direct beneficiary is a developer like Willison — someone who knows how to write a good spec, judge whether the output is sound, and decide it is ready to publish. If you are not a developer, this is not a tool you could pick up this morning and use the same way; the brief, the testing, and the release decision all require technical judgment. What it offers you instead is evidence about where the capability line now sits. A year or two ago, AI coding tools mostly produced snippets that a human assembled. This is a different claim: describe what you want in a few sentences, and the agents handle the building, testing, and packaging end to end.

That matters even if you never write code, because it changes what is plausible in the rest of your life. The gap between "I wish there were a tool that did X" and "a tool exists that does X" is shrinking to the length of a paragraph — at least for the kind of small, well-defined software a library represents.

Now the limits, stated plainly. This is one developer's report about one project, not a benchmark. An alpha release is explicitly unfinished — good enough to publish, not proven. "Very few follow-up prompts" is still not zero, and Willison is unusually skilled at writing the kind of spec that gets good results; a vaguer brief from a less experienced person may fare worse. A library is also a tidy problem with clear success tests; messier software — the kind tangled up with an organization's existing systems — is a harder case this result says nothing about. And while the library itself is available now as open source, the agents involved are commercial tools, so the cost of replicating the experiment is a subscription, not free.

Still, the signal is real. When a cautious, technically credible observer releases what the agents built rather than just blogging about it, that is stronger evidence than another demo video.

productsdeveloperefficiency