AI code generation: what it does well and where it stops
AI code generation is reliable for well-specified work with a clear right answer, and unreliable in proportion to how much judgement a task needs. That single rule predicts most of what you will experience, and it explains why the same tool can feel miraculous on Monday and useless on Tuesday.
This is a practical account rather than a survey: what it handles, where it stops, and how to tell the difference before something ships.
Where it is genuinely strong
- Code with a known shape. A form with validation, a table with sorting, a REST endpoint. These have been written a million times and the model has seen most of them.
- Translation. Between languages, between frameworks, from a data structure to the types describing it. The answer is determined by the input, which is exactly the case models handle best.
- Boilerplate you know how to write and do not want to. Config, scaffolding, the fourth near-identical component. This is the least glamorous use and probably the highest value.
- Getting from nothing to something running. The blank-page problem is real and generation removes it almost entirely.
- Explaining code you did not write. Reading is easier than writing, for models too.
Where it stops
The boundary is not difficulty. Models handle plenty of algorithmically hard problems well. It is whether the task has a right answer discoverable from what the model was given.
- Anything requiring context it does not have. Why your codebase does something unusual, which constraint is load-bearing, what broke last time. It will produce something confident and wrong.
- Decisions with trade-offs rather than answers. Architecture, what to leave out, whether a feature should exist. It will pick one and present it as settled.
- The last ten per cent of a specific change. When the edit is smaller than its description, generation stops paying — this is the point people hit and misdiagnose as the tool being bad.
- Anything where being subtly wrong is expensive. Money, permissions, anything destructive. Not because it fails more there, but because the failures cost more.
The failure mode that actually costs you
Generated code that does not compile is a good outcome — the failure is visible and free. The expensive one is code that compiles, runs, looks right, and quietly does not do what you asked: a filter wired to nothing, a save that saves nothing, a comment where a feature should be.
That is why a compile check is a floor rather than a verdict. It answers whether the code parses, not whether the feature exists, and the gap between those two is where the real cost of generation lives.
Why the same model performs differently in different tools
People compare models and under-weight what surrounds them. A model answering in a chat window is on its own: whatever it produces is what you get. The same model inside a build system gets its output compiled, checked against the original request, and handed back specific instructions when something is missing.
That difference is large enough to swamp a model upgrade, and it grows as the model gets smaller — which is why a modest local model in a good pipeline can beat a much stronger one in a chat box. How the workflow changes locally goes into that trade.
Getting more out of it
- Specify, do not converse. Every ambiguity is a coin flip, and models resolve ambiguity confidently rather than by asking.
- Say what already exists. Naming the file, the function and the pattern you are following removes most of the guesswork.
- Ask for one thing. A request with three parts commonly returns one done well and two stubbed.
- State the constraints that matter, including the ones you think are obvious. Obvious to you is not visible to it.
- Read the dependency list. Every import is code that will run on your machine, and that is a security question rather than a style one.
What it costs, and what drives that
Generation is cheaper than people expect per request and more expensive than they expect per project, because the unit that matters is the iteration rather than the prompt. A build system reads context, generates, reviews its own output and repairs it, so one apparent request can be several calls — and twenty follow-ups is twenty of those.
Which makes the free paths worth knowing about before you commit to a meter. Running a model on your own hardware costs nothing per build no matter how many times you go round, and bringing your own provider key means paying that provider directly with no platform margin stacked on top. What each option actually costs works through the trade.
Reviewing what comes back
Review generated code the way you would review a competent stranger's: assume it is broadly reasonable, and check the places where being wrong is expensive. That is a different activity from reading every line, which nobody sustains.
In practice: does it do all of what was asked, or the easy part of it? Does it handle the input that is not the happy path? Is anything hardcoded that should not be? And does anything touch data, money or permissions — because that is where you slow down and read properly.
The honest summary
AI code generation is a genuine change in how much a small team can produce, and it is not the end of engineering judgement. What it removes is the typing and the blank page. What it does not remove is knowing what should be built, noticing when the result is subtly wrong, and deciding what to do about it.
The people getting the most from it are not the ones who trust it most. They are the ones who know precisely which tasks to hand it, and check the output where being wrong would hurt.