← All posts

Do you need a GPU to run a local coding model?

No — you can run a local coding model on a machine with no dedicated GPU at all. It will work, and on a small model it will be usable. What a GPU buys you is speed, and speed matters more in a build system than it does in a chat window, so the honest answer is that you do not need one and you will notice not having one.

There are two memory numbers, not one

This is the part that confuses people, and everything else follows from it. A model has to be held in memory while it runs, and there are two places that can happen.

  • System RAM, used when the model runs on your CPU. Plentiful and slow.
  • VRAM, the memory on a dedicated graphics card, used when the model runs on the GPU. Scarce and fast.

A 7B coding model at around 4.7GB fits in 8GB of VRAM comfortably and will be quick. The same model in system RAM on a CPU also fits — and generates several times more slowly. Same model, same output quality, very different experience.

Adrian reads your RAM, VRAM and free disk on first run and recommends models accordingly, which is more useful than a specification table because it accounts for what you actually have.

What the no-GPU path actually feels like

Perfectly serviceable for a single component, a utility, a small fix. Tedious for a multi-file feature, because you pay the latency on every pass and building is iterative — twenty turns at forty seconds each is a different afternoon from twenty turns at eight.

It also pushes you toward smaller models than your RAM would allow, which is a real constraint rather than a preference. A 32B model that technically fits your 32GB of system RAM and takes minutes per response is worse to work with than a 7B that answers in seconds. Which model fits your machine covers that trade properly.

Integrated graphics and unified memory

Not every machine splits cleanly into CPU and GPU. Integrated graphics share system RAM rather than having their own pool, so there is no separate VRAM number to fill — and the shared bandwidth is the limit rather than the capacity.

Machines with unified memory sit in a genuinely better position than their specification suggests, because the whole pool is available to the accelerator instead of being capped by a small dedicated allocation. A laptop with a lot of unified memory can run models that would need an expensive discrete card otherwise. Worth knowing before concluding you need new hardware.

Why a GPU helps at all

Worth understanding, because it explains which upgrades are worth money. Generating a token from an LLM means reading through the model's weights, and it happens again for every single token. The work is not especially complicated; there is just an enormous amount of data to move.

So the binding constraint is memory bandwidth rather than raw compute. A graphics card's memory is several times faster to read than system RAM, which is most of where the speed difference comes from. It is also why a faster CPU barely helps: you are not waiting on arithmetic, you are waiting on memory.

The practical consequence is that more RAM lets you run a bigger LLM, and faster memory lets you run it quickly. Those are different purchases solving different problems, and buying the wrong one is the usual expensive mistake.

How much VRAM is enough

The useful way to think about it: the model weights have to fit, with headroom for the context window on top. Every entry in Adrian's catalog carries a download size, and that size is the floor for what the weights need wherever they live.

  • A 7B-class coding model is roughly 4–5GB quantised, so 8GB of VRAM is comfortable.
  • A 14B-class model is around 9GB, which is tight on 12GB and fine on 16GB.
  • A 32B-class model is around 20GB, so 24GB is the realistic floor.
  • Above that you are into workstation and multi-GPU territory, and an API key is usually the cheaper answer.

A model that nearly fits is the worst case. Partial offloading — some layers on the GPU, the rest on the CPU — works, and it is slower than either doing the whole job. Aim for a model that fits with room, not one that only just does.

Before buying anything

Run the free path on what you own first. Local models cost nothing per build, so the cost of finding out is an afternoon rather than a card — and a surprising number of people discover a 7B model inside a build system is enough for the work they actually do.

If it is not enough, the choice is not automatically hardware. A provider key gets you frontier-model quality today at the price you already pay, with no cut taken on top, and what each option actually costs works through when that beats buying a GPU. A card is a fixed cost that pays off only if you build constantly; an API key costs nothing in a quiet month.

The honest summary

A GPU is a speed upgrade, not an entry requirement. Adrian's stated minimum is 8GB of RAM with 16GB recommended for larger local models, and nothing in that mentions a graphics card — because the local path genuinely runs without one.

What a GPU changes is how willing you are to iterate, and iteration is where most of the value in this kind of tool lives. If you find yourself batching up requests to avoid the wait, that is the signal you have outgrown CPU inference — and how the workflow changes locally covers adapting to it before you spend anything.