← All posts

Secure app development when AI writes the code

The security problem with AI-generated code is not that models write exploits. It is that they write ordinary, plausible code that happens to be unsafe — an eval() where arithmetic was needed, a dependency nobody vetted, a manifest with a lifecycle hook in it — and every one of those compiles, runs, and looks correct in a screenshot.

That makes it a different problem from the one most security advice is written for. There is no attacker in the loop and nothing is disguised. The risk arrives as a normal-looking file.

What actually goes wrong

  • Dynamic evaluation. Ask for a calculator and a model will frequently reach for eval() or new Function(), because that is how calculators are written in most of the code it learned from. It works. It also executes whatever ends up in that string.
  • Dependencies nobody chose. Generated code imports packages, and an import is a decision about what runs on your machine. A hallucinated package name that happens to exist on the registry is somebody else's code in your project.
  • Secrets in source. A key pasted into a file to make something work is a key that ends up in your repository, and then in every clone of it.
  • Install-time execution, which is the one people underestimate and the next section is about.

The one that is worse than it sounds

npm runs preinstall, install, postinstall and prepare scripts through the shell during installation. That is a normal part of how the ecosystem works, and legitimate packages depend on it to place platform binaries.

It also means a project manifest containing a postinstall of the shape "download this and run it" is remote code execution triggered by nothing more than opening the project. Not building it, not running it — installing its dependencies. If the manifest arrived from an imported repository, a folder someone sent you, or a model's output, you did not write that line and you will probably never read it.

Adrian strips a project's own lifecycle hooks before running an install on it. Deliberately not with npm's blanket --ignore-scripts flag: that would also suppress dependency scripts, and enough real packages install their platform binary that way that the flag trades a security hole for broken installs. The hooks removed are the ones in the project manifest, which is where injected ones live.

Replacing eval rather than warning about it

When generated code reaches for eval() to evaluate arithmetic, Adrian substitutes a small parser that reads the expression properly — it accepts digits and operators, and returns an error for anything else rather than executing it.

The reason to replace rather than flag is that a warning moves the work to you. A finding you have to read, understand and act on is a finding that gets skipped at 1am, and the failure mode of security tooling is that people turn it off. Fixing it silently and correctly is worth more than telling you about it.

Where your API keys actually sit

If you bring your own key, it is stored encrypted on your machine through the operating system's own credential protection, not in a plain file and not on our servers. Adrian supports keys from OpenAI, Anthropic, Google, OpenRouter, NVIDIA and DeepSeek — you paste one, and your usage is billed by that provider directly with no cut taken.

The desktop build is also hardened at the Electron level: the fuses that would let the packaged application be re-run as a plain Node process, or be attached to with an inspector, are switched off in the shipped binary. That closes a category of local tampering that has nothing to do with what the model wrote.

Running locally changes the exposure, not the code

On local models nothing about your project reaches a model provider, because there is no provider in the loop. For work under a confidentiality obligation that is the difference between a policy question and a non-question.

It is worth being precise about what that does and does not buy you, though. Running locally means your code is not transmitted. It does not make the generated code safer — a local model writes the same eval() as a hosted one. Those are two separate problems and only one of them is solved by where the weights live.

What to check yourself, in order

If you read nothing else in a generated project, read these four things. They are where the real problems cluster, and they take about ten minutes.

  • The dependency list. Every entry is code that will run on your machine. Recognise the names; look up the ones you do not, and be suspicious of a package that sounds right but you have never heard of.
  • Anything handling input from outside the app — a form, a URL parameter, an uploaded file. This is where a mistake becomes someone else's opportunity rather than just a bug.
  • Anywhere a secret appears. A key in source is a key in your git history, and rotating it later is the only fix once it has been pushed.
  • Whatever the app does with a database or the filesystem. Generated code is optimistic about paths and queries in ways that are fine on your laptop and not fine in front of users.

What none of this claims

These are specific defences against specific, common failures. They are not a security review, and passing them means no known unsafe pattern was found — not that the application is secure. Authentication logic, authorisation rules, what you do with user input, and whether the thing should be on the internet at all are decisions no checker makes for you.

The honest framing is that generated code deserves the same review as code from any developer you have not worked with before. What a build system can do is make sure the boring, mechanical, known-bad patterns never reach that review — the same argument as checking that generated code actually does what was asked. Anything that needs judgement still needs yours.