When AI output looks wrong, check your own tools first

An apparent AI failure (invalid output) turned out to be the author's own rendering-tool bug, not the model's fault.


Simon Willison recently spotted what looked like a failure in an AI model's output — a rendering glitch that made the result look wrong — and his first instinct was to blame the model. It wasn't the model. In his own words:

That was entirely incorrect: the rendering glitch was my fault, caused by a bug In my rendering tool . I've now fixed that bug.

The lesson he draws is worth taking seriously by anyone who works with AI assistants: when output looks broken, the model is only one link in the chain, and it is not always the broken one.

What "the pipeline" means for a non-developer

Everything between the AI generating text and you seeing it is a pipeline: the app displaying it, the file format it was saved in, the converter turning it into a document, the clipboard that carried it. Any of those can mangle a perfectly good answer. A missing table might be a spreadsheet import issue. Garbled formatting might be your notes app stripping something it doesn't support. Gibberish in a copied reply might be the copy-paste step, not the model.

Willison's case was a tool he had written himself, which makes the specific bug a developer's problem. But the general habit transfers directly: before you conclude the AI failed, check whether what you're looking at is really the AI's raw output, or the output after something else touched it.

What this looks like in practice

A few cheap checks, before you distrust the assistant:

  • Ask the assistant to repeat or reformat its answer. If it produces clean output the second time, the first display layer is suspect.
  • Look at the same output somewhere else — a different app, a plain text view, a fresh export.
  • If output goes through any tool you configured, automated, or built (a template, a script, a formatting preset), that is where suspicion should start.

Who this is for

If you write your own tools around AI models — as Willison does — this is directly for you: a real case where the bug was in his rendering code, and an honest public correction of it. If you don't write code, the principle still applies, just one level up: the "tool" is whatever app or workflow is showing you the AI's work.

Why it matters

Blaming the model when the fault is elsewhere has a cost: you lose trust in output that was actually fine, you start compensating for a problem that doesn't exist, and — as in Willison's case — you may even publish a wrong conclusion before checking your own side. The reverse failure mode exists too, but the correction here is specific: he asserted something false about a model, investigated, found his own bug, fixed it, and said so.

The honest caveat

This is not a product or a feature — it's a practice, and it's usable today. But it only goes so far: checking your pipeline requires that you can actually see the output before and after your tools touch it. For many people using an AI inside a closed app, that intermediate view isn't available, and the advice reduces to "try another app before giving up." Useful, but not a complete answer to unexplained failures.

productsaccuracyautomationfamily