When an AI agent says it's done and it isn't
The edit is usually fine. The claim about the edit is the problem, and it is a harder one. You ask for a change across four files. The agent works for a few minutes, edits them, and says the change is complete. It is not. Nothing was run, one caller in a fifth file no longer compiles, and you find out in CI or in review or from somebody else. The frustrating part is that the edit was usually reasonable. What was wrong was the sentence at the end.
Why it happens
A language model produces the most plausible continuation of what came before. After a sequence of edits, the most plausible continuation is a summary saying the work is finished, because that is how the thousands of examples it learned from ended. Nothing in that process checks. So "done" is not a claim the model is making about your repository. It is the shape of a closing paragraph. The only thing that turns it into a real claim is a tool that actually ran something, and then reported what came back.
What it costs, and why it is worse than a bad edit
A wrong edit you catch in the diff costs a minute. A wrong claim costs you the assumption you were reviewing under. This compounds. The reason agents are useful is that you stop reading every line and start reading the summary. Once the summary is unreliable you have to go back to reading every line, and at that point the agent has moved the work rather than done it.
How to tell the difference
The question to ask of any tool is simple: did it run anything, and can you see what came back?
- Did it execute your tests, or describe executing them? Those look similar in a summary and are not the same event.
- Is the raw output on screen? A tool that ran your suite has output. A tool that did not will paraphrase.
- What does it do when it cannot verify? The honest behaviour is to say so. A tool without that check tends to report success anyway.
- Did it see the file it broke? An agent reading only your open tabs cannot verify a change it cannot observe.
A test you can run in ten minutes
Take a repository you know well with a test suite that passes. Ask the agent for a change touching at least three files. Then, before you look at the diff, break something it just wrote by hand and ask it to continue. A tool that runs your suite notices. A tool that does not will keep going and tell you everything is fine. That single exercise tells you more than any comparison page, including ours.
What to do if your current tool does this
- Make the tests the specification. An agent iterating against a suite is only as good as the suite. Break the behaviour a test guards and confirm the test goes red, because a test that passes either way actively steers an agent wrong.
- Shrink the unit of change. Review fatigue is what turns an unverified claim into a merged defect.
- Ask for the commands. If the tool can show what it ran, make it. If it cannot, treat every summary as a draft.
Where AstraCode fits
This is the problem we built AstraCode around. It plans the change, makes it, runs your tests, reads the output, fixes what it broke, and runs them again, with what it ran on screen rather than summarised. When it cannot verify something it says so instead of reporting success. There is a free tier with no card, and the ten-minute exercise above is a fair way to judge it, or any other tool.
Originally published on the AstraCode blog.
Comments
No comments yet. Start the discussion.