Qwen 3.8 27B defaults to a maximum reasoning-effort setting that causes it to wildly over-think even trivial requests, so you should run it at low or no reasoning first.
When Simon Willison ran the newly released Qwen 3.8 27B model on his own machine, he discovered that it ships with its reasoning-effort dial turned all the way up by default. The result: the model spent 21 minutes thinking through a question that did not deserve it. His verdict was blunt — the default is not how anyone should run the model, especially on ordinary consumer hardware.
"Reasoning effort" is a setting many modern AI models expose that controls how much internal deliberation the model does before it answers. Turned up high, the model works through problems step by step — useful for genuinely hard questions in math, logic, or planning. Turned down low or off, it just answers. The catch is that all that deliberation takes time and computing power. On a cloud service you may barely notice the delay; on your own laptop or desktop, a maximum-effort model can turn a simple question into a long wait for an elaborate answer you never asked for.
Willison's recommendation is to treat the shipped default as a mistake and start at the other end of the dial:
"My strong recommendation: ignore that default. Run Qwen 3.8 27B on low or even no reasoning levels at first. It's a great model, but wow that default setting is a bad place to start."
He was harsher still about the setting itself:
"This is a hilarious default. It's absolutely not a good way to run the model, especially on consumer hardware."
This matters most to the growing number of people who run AI models locally — on their own hardware rather than through a subscription service like ChatGPT or Claude. People do this for privacy, for cost, or because they like controlling their own tools. If that is you, and your local model seems slow or produces sprawling, over-engineered answers to simple requests, the reasoning-effort setting is the first thing to check. Turning it down is free, takes seconds, and may transform how usable the model feels. The practical order of operations: start at low or no reasoning, and only reach for higher effort when a task actually stalls or comes back wrong.
If you do not run models locally, this mostly does not apply to you. Hosted services choose these settings for you. But the underlying lesson travels: when an AI tool misbehaves, the fix is often a setting, not a different tool.
This is usable now. Qwen 3.8 27B is shipping, and reasoning-effort controls exist in the tools people use to run local models today. It is not a proposal or a research idea — it is a configuration note from someone who ran the model and timed the result.
The honest limits: a low-reasoning model is faster but shallower. For a genuinely difficult problem — a tricky bit of analysis, a multi-step plan — you may need to turn the dial back up and accept the wait. The setting is a trade-off, not a free lunch. And Willison's report is one person's experience with one model on his hardware; how much the default hurts you will depend on your machine and what you ask. But the asymmetry is the point. A model that over-thinks easy questions wastes your time constantly; a model that under-thinks a hard one fails visibly, and you can simply ask again with more effort. Starting low costs you little. Starting at maximum, as the default does, cost Willison this:
"Was that worth waiting 21 minutes for? Absolutely not."