AI models may use simplified, grammatically imperfect English in their internal reasoning traces to save tokens and increase efficiency.
If you have ever peeked under the hood of an AI assistant while it "thinks," you may have noticed something odd: the reasoning it shows you does not read like the polished answer it eventually gives you. Simon Willison, a widely followed writer on AI tools, noticed this too while watching a model's reasoning trace — the stream of text some assistants produce as they work through a problem before answering.
"It's interesting how the reasoning trace uses slightly truncated English, presumably because perfect grammar isn't useful or token efficient for hidden reasoning text."
His observation is small but revealing. The behind-the-scenes text tends toward shorthand — clipped sentences, dropped words, compressed grammar — because the model has no reason to write beautifully for an audience. Each word it generates costs tokens, the small units of text that models process, and generating fluent prose is more expensive, in a loose sense, than generating fragments. For text nobody is meant to read, fragments are enough.
What this means for you depends on how you use these tools. Most assistants let you expand a "thinking" panel to watch the model reason — a feature shipped in several popular chatbots that shows the intermediate steps before the final answer. If you have opened one and found the text strange, terse, or oddly ungrammatical, you were not watching a malfunction. You were watching the difference between a draft and a deliverable. The reasoning trace is scaffolding; the answer is the building.
There is a practical reason to know this. Some users — often people evaluating answers in fields like law, medicine, or research — read reasoning traces to check how an assistant arrived at a conclusion. If you are one of them, the shorthand style is worth understanding rather than dismissing. A truncated sentence in the trace is not necessarily a shallow thought; it may just be an efficient one. Judging the reasoning by its polish is a mistake, the same way judging a person's private notes by their penmanship would be.
At the same time, a limit worth stating plainly: a reasoning trace is not a transcript of what the model "really" did. It is text the model generated, in the same way it generates everything else. The shorthand style makes it look candid and unfiltered, but there is no guarantee it faithfully represents the internal computation. Treat it as a rough sketch, useful for sanity-checking, not as evidence of the model's inner workings.
It is also worth being honest about who this matters to. If you simply use an assistant to draft emails or answer questions and never open the thinking panel, this observation changes nothing about your day. The polished output is the product; the trace is incidental. This is material for a narrower group: people who read reasoning traces deliberately — curious users auditing how an answer was produced, researchers studying model behavior, and developers debugging why a model went wrong. For that audience, Willison's note is a useful calibration: the rough grammar is a feature of the format, not a bug or a red flag.
The observation itself is available and verifiable today — it is not a proposal or a rumor. Any user can open a reasoning trace in a chatbot that exposes one and see the same truncated style for themselves. What remains open is the "presumably" in Willison's phrasing: the explanation that it saves tokens is an inference, not something confirmed by the labs building the models. The style is observable; the motive is a reasonable guess.
The broader takeaway is a small piece of literacy for anyone working alongside AI: the assistant's working notes do not have to look like its finished work, and the gap between the two tells you something about how these systems allocate effort — fluency where it counts for the reader, economy everywhere else.