Unguided LLMs produce consistent findings only 50% of the time when run repeatedly on the exact same input and prompt.
Run an AI assistant on the same task five times and you may get five different answers. Not slightly different phrasing — actually different findings. According to Manoj, an assistant asked to analyse the same codebase with the same prompt five times in a row produced the same set of findings only about half the time.
If you run it against the same repo, same code, same prompt five times, only 50% of the findings are consistent across the runs.
This is not a bug to be fixed. It is how these systems work at a basic level: they generate responses by picking likely next words, and that process involves a degree of chance. Two runs that start identically can diverge early and end up in different places. The practical consequence is that the number above is a ceiling for an unguided model — a raw assistant, given a task and left to it, is roughly a coin flip for repeatability on analytical work.
Why that matters depends on what you are asking it to do. For a one-off summary or a draft you will edit anyway, inconsistency is a nuisance at most. The number becomes a real problem when you delegate a repetitive analytical task — reviewing contracts, triaging reports, evaluating candidates against a rubric, auditing expenses — where the whole point is that the assistant applies the same judgement every time. If half the output shifts between runs, you cannot tell whether a change in results reflects a change in the input or just the roll of the dice. You end up re-checking the assistant's work, which was the labour you were trying to save.
The people this most directly concerns are those building or buying evaluation workflows — systems where an assistant scores, flags, or filters things at volume. The honest version of the audience here skews technical: the specific figure Manoj cites comes from running a model against a code repository, which is developer work, and the teams who will act on it first are engineering and operations teams standing up automated review pipelines. If you are a non-developer who simply uses an assistant day to day, the takeaway is narrower but still useful: do not treat a single run as a verdict. If an answer matters, ask again, and be suspicious of any automated process that nobody spot-checks.
The finding is also an argument for the current direction of the field. The reason "guardrails" and "structured context" have become industry vocabulary — checklists the model must follow, fixed output formats, examples of correct answers baked into the prompt, explicit criteria rather than open questions — is precisely this 50% figure. Structure narrows the space of acceptable answers, which narrows the variance. None of that eliminates the underlying randomness; it fences it in.
This is usable knowledge today, not a prediction. The behaviour Manoj describes is a measured property of shipping systems, not a limitation scheduled for a future release. What is not resolved is the harder question underneath: 50% consistency is a number about variance, not accuracy. A perfectly consistent assistant could be consistently wrong, and the figure says nothing about how often the findings it does produce are correct — that is a separate measurement nobody is quoting here.