Local LLM development: how the workflow changes
Local LLM development works when you stop treating the model as a smaller version of a hosted one and change the shape of the work instead: smaller steps, tighter scope per request, and more structure around the model rather than more capability inside it. People who try to use a 7B model the way they use a frontier one conclude local models are not ready. People who adjust the loop mostly do not.
The difference is context, not intelligence
The gap between a local model and a frontier one shows up least in single, well-specified tasks and most in long ones. Ask for a component with clear requirements and a good 7B coding model does well. Ask it to hold six files, an architecture decision and a half-finished refactor in its head, and it does not.
So the workflow adjustment is mostly about how much you ask for at once. A request that a frontier model would handle in one pass is often three requests locally, and the three-request version frequently produces better code, because each one was specific enough to get right.
Work in smaller units
- One concern per request. "Add the filter" beats "add the filter, wire it to the header, and fix the empty state", which locally tends to produce one of the three done well and the others stubbed.
- Say what already exists. A local model benefits more than a hosted one from being told which file and which function, because it has less room to work it out.
- Check after each step rather than at the end. A wrong assumption compounds, and finding it three steps later costs all three.
- Prefer editing over regenerating. Regeneration re-rolls decisions that were already correct.
Structure does the work capability would have
This is the part that decides whether local development is viable, and it is why the same model performs so differently in different tools. A model answering in a chat box is on its own. A model inside a build system gets its output compiled, checked against the request, and handed back specific instructions when something is missing.
That loop matters much more at 7B than at frontier scale, because there is more for it to catch. The checks that run after the compile gate are the difference between a small model producing a stub you have to find and a small model producing a stub that gets repaired before you see it.
The practical consequence is that comparing local and hosted models by chatting to both is a misleading test. It measures the model alone, which is not how either will be used.
Latency changes what you ask for
A local build is slower per iteration, and this shapes behaviour more than people expect. When a response takes forty seconds rather than four, the cheap habit of firing off a half-formed request and seeing what comes back stops being cheap.
That is not entirely a loss. Being made to write a precise request is a discipline that produces better output from any model, and plenty of people find their prompts improve on local hardware for exactly this reason. But it is a real cost on a thin laptop, and which model fits your machine is worth reading before concluding local development is slow — a model that just fits and one with headroom behave very differently.
Budget for two models, not one
Coding and visual design are different skills and the models good at them are different models, so a working local setup usually means two sets of weights resident rather than one. That is a memory question before it is a preference: the pair has to fit alongside your editor, your browser and the operating system.
It is also a place people over-buy. The design pass does not need your largest model, and pairing a strong coder with a small designer is usually a better use of the same memory than one big model doing both jobs adequately.
When to reach for a hosted model anyway
- The first build of something substantial, where there is no existing code to anchor the work and the scope is inherently large.
- A cross-cutting change touching many files at once, which is the case local models handle worst.
- Anything where you cannot describe what you want precisely. Vagueness is exactly what frontier capability absorbs.
- When you are debugging the model rather than the app. If three attempts have not landed, the fourth usually will not either.
Switching is per role rather than all-or-nothing — you can point the coder at a key and leave the design pass local, or the reverse. Local stays available alongside it, so this is a dial rather than a decision.
What the trade actually buys
Zero marginal cost per build, which changes behaviour on its own — nobody rations iterations against a meter that reads nothing. Work that stays on your disk. And a setup that keeps working with the network unplugged, which is the whole argument for running the builder locally if you are ever on a locked-down network or a plane.
The honest version
A local model inside a good pipeline is genuinely useful and is not secretly equal to a frontier model. Anyone claiming otherwise is selling something. What is true is that the gap is much smaller than a chat-box comparison suggests, and that for a lot of ordinary work — a component, a utility, a fix, an iteration on something that exists — the difference stops mattering.
The useful test is not whether local matches hosted. It is whether local is good enough for the work you actually do, at a price of zero, and for a surprising amount of that work it is.