← All posts

Reproducible builds when the code was generated

Run the same prompt twice and you get different code. That sounds like it should be a problem for reproducible builds, and it mostly is not — because the thing worth reproducing was never the generation step. It is the build: the same source producing the same artefact, every time, on any machine.

Separating those two is what makes the question tractable, and conflating them is why the topic sounds harder than it is.

Two different reproducibility questions

The first is whether a prompt reliably produces the same code. It does not, and largely cannot: sampling is probabilistic, and even at the most deterministic settings, a model version change or different hardware moves the output. Treating a prompt as a build input is a losing position.

The second is whether committed source reliably produces the same running application. That one is an ordinary engineering problem with well-understood answers, and it does not care in the slightest how the source came to exist.

Once code is generated it is just code. The prompt is history, like the conversation in which a feature was specified — worth keeping, never a substitute for the artefact.

The code is the source of truth, not the prompt

This has one practical consequence that matters more than the rest of the article: generated code belongs in version control, immediately, exactly like code you typed.

The temptation is to think of the prompt as the real input and the code as an output that can be re-derived. It cannot. Regenerating gives you different code with different bugs, so a project whose source of truth is a prompt has no source of truth at all. Commit the code; keep the prompt in the commit message if it is useful.

Owning the output is what makes this possible in the first place. If a tool keeps the result inside its own platform, the question is moot — you cannot version what you cannot hold, which is one reason where the code actually lives is worth checking before committing to any builder.

What actually breaks reproducibility

  • Unpinned dependencies. A caret range means "whatever was newest when you installed", so the same source resolves differently in March and in June. This is the big one and it has nothing to do with AI.
  • A missing lockfile. The lockfile is the record of what was actually resolved; without it the dependency list is a preference rather than a fact.
  • Install-time scripts. Packages that run code during installation can behave differently per machine, which is also a security question — a manifest hook is remote code execution triggered by installing.
  • Environment assumptions. A path, a locale, a Node version. Generated code is optimistic about all three, because it was written on nobody's machine in particular.

Notice that generation makes none of these worse. It does make the last one more likely to go unnoticed, because nobody sat there choosing the assumption.

Verifying a build you did not compile

There is a related problem for anyone downloading software rather than building it: how do you know the binary you have is the one that was published?

The ordinary answer is a published checksum. Every Adrian release ships with a SHA-256 you can compute against the file you downloaded — if the two match, the bytes are the ones that were released, and if they do not, something is wrong regardless of where the file came from. The download page has the current one and the command to check it.

Worth being honest about the limits of that. A checksum proves the file matches what was published; it does not prove the published file is trustworthy. It defends against a corrupted download or a substituted mirror, not against the publisher. That is what code signing addresses, and the current Windows build is unsigned — which is why it triggers a SmartScreen warning on first run.

The determinism you can actually get

Generation is not repeatable, but the checks around it are, and that turns out to be the more useful property. A compile gate gives the same verdict on the same source every time. So does a check that reads the code back against the request. Those are deterministic even when the thing they inspect was not.

Which reframes the goal. You are not trying to make the model produce identical output; you are trying to guarantee that whatever it produced meets a fixed, repeatable standard before you see it. That is a bar generation can be held to, and it is why checks that run after the compile gate matter more than sampling settings.

A practical setup

  • Commit the lockfile and pin versions exactly. Every reproducibility conversation ends up here.
  • Commit generated code the moment it works, before you iterate on it. The version that worked is the one worth being able to return to.
  • Record what produced it — model, rough prompt, date — in the commit message. Useless for regenerating, useful for understanding a decision six months later.
  • Verify downloads by checksum, including your own releases. It costs one command and catches the whole class of corrupted-artefact problems.

The honest summary

AI generation does not make builds less reproducible, because it operates at a stage before the build begins. What it does is make it tempting to treat code as disposable and regenerable, and that temptation is the actual risk — not the sampling temperature.

The discipline is unchanged and slightly more important: commit the artefact, pin the inputs, verify what you ship. The same rules that applied when a person typed it.