Running a strong AI model locally on your own machine

With 32 GB of RAM or more, you can run Muse Glimmer locally (e.g., via LM Studio's 18.16 GB version) and still have room to run other applications at the same time.


Simon Willison recently ran a model called Muse Glimmer entirely on his own computer — not through a website or an API, but locally, using a packaged version distributed through LM Studio. He showed off the result plainly: a pelican image, generated by the model, right there on his machine.

Here's a pelican which I generated using LM Studio's 18.16 GB version of the model

The claim underneath the pelican is the interesting part. Local AI models — ones that run on your hardware rather than on a company's servers — have a reputation for needing a lot of memory. RAM is the constraint: the model has to live in it while it runs, and the bigger the model, generally the more capable it is. This version of Muse Glimmer takes 18.16 GB, which sounds enormous until you hear Willison's reasoning:

I really like this size of model, because if a machine has 32 GB of RAM or more (mine has 128GB) it leaves plenty of space for running other applications at the same time.

In plain terms: if your computer has 32 GB of memory, this model occupies a bit more than half, and your browser, documents, and everything else still fit alongside it. You are not dedicating a machine to the model. It sits on an ordinary desktop or laptop and shares.

Why would anyone bother, when chatbots on the web are free or nearly so? Two reasons, mostly. The first is privacy. Everything you type into a hosted assistant travels to someone else's infrastructure. A local model never leaves your machine — nothing is logged, retained, or used for training by a provider, because there is no provider. The second is independence. There is no subscription to lapse, no rate limit, no outage, no company that can change the model's behavior or retire it. If the file is on your disk, it keeps working.

Who is this actually for? This one honestly lands with a technical-ish reader. Running a local model means installing software like LM Studio and choosing a model variant, and the audience Willison is writing for — people who benchmark models and generate test images of pelicans — skews developer-adjacent. That said, the barrier here is lower than the phrase "run a model locally" suggests. LM Studio is a desktop application, not a command-line exercise; if you can install an app and download a large file, you are most of the way there. The reader who benefits most is someone privacy-conscious enough to want AI that doesn't phone home, on hardware they already own, rather than a developer doing anything clever with it.

Is it real, or just an idea? It's shipping. The model exists, the 18.16 GB package exists, and Willison ran it and published the output. This is not a roadmap.

Now the limits, which a vendor's pitch would skip. Eighteen gigabytes is a large download, and 32 GB of RAM rules out a lot of perfectly good laptops — many ship with 8 or 16. A model this size is capable for its class, but "capable" is not "frontier": it will not match the largest hosted models on hard problems, and the brief doesn't include any benchmark numbers to argue otherwise — Willison's evidence is a pelican, not a scoreboard. Whether Muse Glimmer costs anything, and under what license, isn't stated here either. And "plenty of space for other applications" is true in memory terms, but a model generating text will still make your machine work hard while it does.

The honest summary: if you already have a machine with 32 GB of RAM and a reason to keep your prompts private, a local model of this size is a practical thing today, not a hobbyist stunt. If you have 16 GB and no privacy requirement, the hosted tools remain the easier answer.

privacyfinancehomeproducts