← All posts

Why your local LLM is slow, and what to check first

A local LLM that has gone from usable to unbearable is almost always running out of memory rather than running out of compute. When the weights plus the context no longer fit, the machine starts moving data to disk to cope, and disk is thousands of times slower than RAM. That is not a gradual slowdown; it is a cliff.

Which is why "buy a faster CPU" is usually the wrong response. Four things account for nearly all of it, and they are worth checking in this order.

1. The model does not actually fit

The tell is distinctive: not merely slow but pathologically slow, with heavy disk activity, and the whole machine turning sluggish rather than just the model. If everything else on your computer also got worse, this is what is happening.

The fix is to drop to a smaller model, not to wait it out. Adrian's catalog carries a minimum RAM figure per entry for exactly this reason, and that figure allows for the operating system and your other applications — a model whose weights technically fit your total RAM does not fit the RAM you actually have free. Which model fits your machine goes through the tiers.

2. The context window grew

The most common cause of a model that was fine yesterday and is slow today, and the least obvious. Context is held in memory alongside the weights, and it grows as the conversation or the project grows. A model comfortable on a small project can be at its limit on a large one, without you changing anything.

It also explains why this bites harder in a build system than in a chat. An agent that reads several files, plans, generates, then re-reads its own output to check it fills a context window far faster than someone asking a question does. Catalog context windows run from 32K to 1000K tokens, and the larger the window you actually use, the more memory you need on top of the weights.

Practical response: start a fresh session rather than continuing an enormous one, and keep requests scoped to the files that matter instead of the whole project.

3. It is running on the CPU when you expected the GPU

If a model is a little too large for your VRAM, some layers end up on the graphics card and the rest on the processor. That works and it is slower than either alone, because every token now waits on the slow half.

A model that fits your VRAM with headroom will outrun a larger one that spills, often by a wide margin — which is the opposite of what the parameter counts suggest. Whether you need a GPU at all covers the memory arithmetic behind that.

4. The quantisation is heavier than you think

Models in the catalog ship quantised — stored at reduced precision so they take less memory. Most are Q4, some Q4_K_M, and GPT-OSS 20B is MXFP4. This is what makes local inference possible at all, and it is also a lever: a more heavily compressed build of a larger model is not automatically better than a lighter build of a smaller one.

So if you have picked the biggest model your machine will hold, you may have optimised the wrong number. The size that fits comfortably usually beats the size that fits at all.

You are probably holding two models, not one

Easy to forget when diagnosing memory. Coding and visual design are different skills and Adrian assigns them to different models, so a build can have a coder and a designer resident at once. Add them up before concluding a single figure fits.

This is also where the cheapest speed win usually hides. The design pass does not need your largest model, and pairing a strong coder with a small designer frees real memory for the half of the job you care about — often more than switching the coder would.

Measure before changing anything

One pass with a system monitor open answers most of this. Watch memory while a build runs: if it climbs to the ceiling and disk activity spikes, it is cause one or two. If memory sits comfortably and it is still slow, it is three or four, and swapping models is pointless.

Thirty seconds of looking beats an afternoon of guessing, and it stops you buying the wrong upgrade — which is the expensive version of this mistake.

What is usually not the problem

  • Your CPU. Token generation waits on memory bandwidth, not arithmetic, so a faster processor barely moves the number.
  • Your internet connection. Nothing leaves the machine on the local path — if speed varies with your network, something is not running locally.
  • The first response after a cold start. Loading weights into memory takes a moment; judge the second request, not the first.
  • Adrian itself. The pipeline adds planning, checking and repair passes, which are more calls rather than slower ones — real cost, but visible as extra steps rather than sluggishness.

When it is as fast as it is going to get

There is a point where the hardware is the hardware, and the answer is to change how you work rather than what you run. Smaller requests, tighter scope, editing rather than regenerating — how the workflow changes locally is about exactly that adaptation.

And if the loop is still too slow to be worth it, the other free route is a provider key: you pay the provider directly at their rates with no cut taken, and get frontier speed without buying hardware. What each option costs covers when that trade makes sense — which for a thin laptop is often immediately.