← All posts

Natural language to code: where the meaning gets lost

Natural language to code works well until your description runs out, and then it keeps going anyway. Everything you did not specify still has to become something, so the model picks — confidently, invisibly, and without telling you which choices it made. That is the whole difficulty, and it is a property of the translation rather than a defect in any particular model.

Every sentence is underspecified

Take "add a search box that filters the list". Perfectly clear to a person. Now count the decisions it does not contain: does it match on every field or one? Is it case-sensitive? Does it filter as you type or on enter? What shows when nothing matches? Does it survive a page refresh?

A human developer resolves these by asking, or by knowing the product well enough to guess correctly. A model resolves them by producing the most statistically ordinary answer for each, independently. Most of the time that is fine. When it is not, the wrongness is buried inside something that otherwise looks exactly right.

The gaps are filled silently

This is the part that makes the failure expensive rather than merely annoying. A compiler tells you what it could not resolve. A model resolves everything and reports nothing, so there is no list of assumptions to check — the assumptions are simply in the code now, indistinguishable from the parts you actually asked for.

It is also why "it did not do what I asked" is usually inaccurate. It did what you said. The gap is between what you said and what you meant, and that gap is invisible until you use the thing.

Writing a description that survives translation

  • Say what happens when it goes wrong. Empty states, no results, bad input. Models default to the happy path, and the happy path is rarely where a feature disappoints.
  • Name the thing you are changing. "In the header component" removes an entire class of confident wrong answers.
  • State the constraint you think is obvious. Obvious to you means invisible to it.
  • Give one example of correct behaviour. A concrete case resolves more ambiguity per word than any amount of description.
  • Ask for one thing. Three requests in a sentence usually returns one done and two stubbed.

None of this is about prompt tricks or magic phrasing. It is closer to writing a good bug report: specific, bounded, and explicit about the conditions.

The better fix is being asked

Writing perfect descriptions is a skill, and expecting it of everyone is unrealistic — most people do not know which details matter until something is wrong. The more useful fix is for a vague request to produce a question rather than a guess.

That is why Adrian turns a prompt into a build plan before any code is generated, and vague prompts get an interview step instead of an assumption. A wrong assumption at the start is the most expensive kind: every later stage builds on it, and you find out at the preview, having paid for the whole build. The full pipeline walks through where that sits.

Translation errors need catching, not just avoiding

Better prompts reduce the gap and do not close it. Something still has to check that the code matches the request, which is a different question from whether the code is valid.

A compile check confirms the code parses. It cannot confirm the search box filters anything — a filter wired to nothing compiles perfectly. Checks that read the generated code back against what was asked are what catch a translation error, and why generated apps compile but do not work is the longer version of that argument.

Why the second attempt usually lands

In practice most people do not write a precise description first. They write a rough one, look at what comes back, and correct it — and that works, because seeing a wrong answer is a very efficient way to discover which detail you left out.

That makes iteration cost the thing that matters most. If each round is free, being vague first is a reasonable strategy. If each round is metered, precision is worth investing in up front. It is one of the better arguments for starting on a local model, where the loop costs nothing.

Some things are faster to point at than describe

There is a category of change where language is simply the wrong instrument. Spacing that is slightly off, a colour half a shade wrong, this element aligned with that one — describing those precisely takes longer than making them, and the description is more likely to be misread.

Which is why an editor beside the preview matters more than it sounds. The realistic workflow is not pure prose; it is describing what is easy to describe and directly editing what is not. A tool that only accepts descriptions turns the last ten per cent of a project into a negotiation.

The honest limit

Natural language will not become a precise specification language, because the ambiguity is the point of it — that is what makes it fast to write. Code is unambiguous because it must be. Translating between them will always involve someone or something deciding what you meant.

So the realistic goal is not to eliminate the gap but to make it cheap: get asked when it matters, get told when the result does not match the request, and keep the loop fast enough that being wrong once costs an afternoon rather than a week.