AI code refactoring: changing code that already works
Refactoring is harder for a model than writing new code, and the reason is that success is defined by what does not change. New code only has to satisfy your description. A refactor has to satisfy your description and preserve every behaviour nobody wrote down — including the ones that exist by accident and that something else quietly depends on.
The bar nobody states
When you ask for a component to be split up, you are also asking for a hundred things you did not mention: the same props, the same render order, the same behaviour when the list is empty, the same thing on a slow network. None of that is in the request. All of it is in the requirement.
A model cannot tell which of those behaviours are load-bearing and which are incidental, because that information is not in the code. It is in your head, or in a bug report from eight months ago. So it preserves what looks intentional and normalises what looks accidental — and sometimes the accident was the fix.
Editing and regenerating are not the same operation
This is the distinction that decides whether a tool is usable on a real project. Regenerating produces a fresh answer to your description. Editing changes the specific thing you asked about and leaves everything else alone.
A system that regenerates on every prompt will cheerfully undo the fix it made two prompts ago, because that fix was never in your description — it was a correction you made along the way. You end up re-reporting the same bug, which is the experience people describe as fighting the tool. How a second prompt should behave is the difference between a demo and something you can use past day one.
What refactors well
- Mechanical, local changes. Renaming consistently, extracting a function, converting a callback chain — the answer is determined by the input and there is little to get wrong.
- Repetitive changes across many files. Tedious for a person, and tedium is exactly where human attention fails and a model's does not.
- Translation work. Between frameworks, into types, out of a deprecated API. There is a right answer and it is derivable.
- Cleanups you can verify at a glance. If reviewing it is fast, the risk is low regardless of how the change was produced.
What does not
- Anything with subtle behaviour you cannot restate. If you cannot describe what must stay true, nothing can check that it did.
- Code you do not understand. A refactor you cannot review is a rewrite you are accepting on trust, and the whole point of refactoring is that behaviour is preserved.
- Architectural change. Moving from one pattern to another is a series of trade-off decisions, and models settle trade-offs by picking the most common option rather than the right one.
- The load-bearing weird part. That comment saying "do not remove, breaks Safari" is exactly what a tidy-up will remove.
Tests are what make this safe
The traditional advice is that you should not refactor without tests, and AI does not soften that — it sharpens it. Tests are the only mechanism that states the behaviour a refactor must preserve, and without them you are comparing the new code against your memory of the old.
Which suggests an order most people get backwards. If a piece of code needs refactoring and has no tests, the first thing to generate is the tests — against the current behaviour, before touching it. That is also a task models are good at, because the specification is right there in the existing code.
It has to read before it can change
New code can be written from a description alone. A refactor cannot — the tool has to find the relevant files first, and how well it does that decides the quality of everything after. A model that rewrites a file it never read is how a follow-up quietly undoes the previous one.
This is also why refactoring costs more per request than generating on a paid model: the existing code goes into the prompt every time, so the bill scales with the project rather than the change. On a local model it costs nothing but time, which is one of the better arguments for keeping the loop on your own machine while iterating.
It is worth being blunt about the limit here. Beyond a certain project size, no tool can put everything relevant in front of the model at once, so it works from a selection. Refactors that depend on a relationship the selection missed will be wrong in ways that look confident, and no amount of prompting fixes a file the model never saw.
Reviewing a refactor
Read the diff, not the file. A refactor's whole claim is that behaviour is unchanged, so what matters is precisely what moved — and a rewritten file hides that by presenting everything as new.
Be most suspicious of the parts that got shorter. Code shrinks honestly when duplication goes, and dishonestly when a condition someone added for a reason quietly disappears. A refactor that removed a branch is worth understanding before accepting.
How to ask for one
- Name the file and the function. Scope is the main control you have over a refactor's blast radius.
- Say what must not change, especially the odd bits. "Keep the existing empty-list behaviour" costs six words and saves a bug.
- Ask for one transformation at a time. Combined refactors produce diffs nobody can review, and an unreviewable diff defeats the purpose.
- Refactor and add behaviour in separate steps, never together. Otherwise you cannot tell which change broke what.
The last point is old advice that predates any of this, and it matters more now, not less — precisely because generating a combined change has become so easy.