The Agent Shouldn't Be Able to Approve Its Own Rules
Once a coding agent can change the architecture, changing the rules that protect it becomes a different kind of operation. One of the less obvious problems I've run into with coding agents isn't that they make bad changes. It's that sometimes they make a perfectly reasonable change that invalidates one of the rules I'm using to check the project. That's a different problem. Imagine a project has this rule: All payment-provider access must go through PaymentAdapter. An agent is asked to add support for another payment provider. It looks at the existing code and decides that the current adapter isn't quite right. It wants to introduce a new abstraction. The resulting change crosses a boundary that the project currently protects. The checker reports a violation. So what should happen next? The obvious automation is: - agent changes the code; - check fails; - agent changes the rule; - check passes. Technically, everything worked. But the system has a problem: The thing being checked was allowed to change the conditions of the check. That made me much more interested in the difference between changing the code and changing the rules around the code. The dangerous loop The simplest version looks like this: agent ↓ change code ↓ verification ↓ failure ↓ agent changes policy ↓ verification ↓ pass There is nothing obviously broken here. The agent might even have a good reason for changing the policy. The problem is that the verification boundary has disappeared. A policy is supposed to tell the system what has to remain true. If the same actor that made the change can also redefine what "true" means, a successful verification doesn't tell you very much. It's just a moving target. And this isn't specific to architecture. You could do the same thing with: maximum service size = 300 lines The agent produces a 420-line service. The check fails. The agent changes the limit to 500. The check passes. Or: module A cannot import module B The agent needs the dependency. It changes the rule. The import is now allowed. Again, maybe that's the right architectural decision. But those are two separate decisions: - The implementation changed. - The policy changed. They shouldn't become one operation just because the same agent proposed both. Not every policy violation is actually a mistake This is where it gets more interesting. I don't want an architecture checker that treats every violation as proof that the agent did something wrong. Sometimes the agent really should change the architecture. Suppose an application has: Checkout ↓ PaymentAdapter ↓ Stripe And I decide that the product is getting large enough that payment workflows deserve their own domain boundary: Checkout ↓ PaymentService ↓ PaymentAdapter ↓ Stripe The change introduces new files. Some imports move. Some old boundaries disappear. New ones appear. A strict checker could report a pile of violations. That doesn't mean the change is bad. It means the current policy describes the old architecture. This distinction matters. A policy isn't supposed to prevent architecture from ever changing. It is supposed to make architecture changes explicit. Proposal and approval are different things This led me to a fairly simple rule: An agent should be able to propose a policy change, but it shouldn't automatically be able to approve that policy change. For example: Current policy: Payment provider access must go through PaymentAdapter. The agent can say: Proposed change: Allow PaymentService to depend directly on a new internal PaymentProvider interface. Reason: The current adapter boundary prevents the new workflow from sharing transaction state correctly. That's useful. The agent has done the hard reasoning. It has identified the existing constraint. It has explained why the constraint may no longer fit. But the proposal should remain a proposal. The important part is that the authority approving the policy change is separate from the agent that authored the change. Otherwise the system can silently move the goalposts. This is the same reason I don't want approval in the prompt A prompt can say: Never deploy without approval. That's useful instruction. It's not much of a control if the same process can modify the configuration that defines what counts as an approved deployment. The more autonomous the agent becomes, the more these distinctions move out of the prompt and into the environment around it. The agent should be able to reason about the policy. It should be able to request a policy change. It should be able to explain the change. The actual enforcement shouldn't depend on the agent remembering to follow its own instructions. A policy change should leave a trail Once policy becomes a real project artifact, another problem appears. You need to know what the policy was when the original change was checked. Otherwise you can end up with a strange situation where today's successful verification only makes sense because yesterday's policy was replaced. That's why I like keeping policy changes explicit and versioned. Something like: project state ↓ policy revision 17 ↓ agent proposes architecture change ↓ verification fails under policy revision 17 ↓ policy proposal ↓ approval ↓ policy revision 18 ↓ verify change again Now there are two separate facts: - The change passed under policy revision 18. - Policy revision 18 was itself approved. That's much more useful than simply seeing a green check. The old policy still matters There's another subtle point here. Suppose an agent changes the rule and then verifies the same diff against the new rule. You can no longer tell whether the original change violated the previous policy. So I want the system to preserve the distinction between: what the project allowed before the change and: what the project allows after the change This becomes especially useful when investigating a change later. You can ask: - Why did this dependency become allowed? - Was it always allowed? - Was there an explicit exception? - Did a policy revision happen at the same time? - Who approved it? - What evidence led to the change? Those questions are much harder to answer when policy is just another mutable config file. Temporary exceptions are different again Sometimes the policy is fine. The violation is temporary. For example: All persistence access must go through Repository. But I'm in the middle of a migration. I don't want to remove the rule. I just need one known exception for two weeks. That's not really a policy change. It's a waiver. And I think treating it as a different object makes the whole system easier to reason about. A useful waiver has at least: owner reason scope expiry So instead of: remove the rule you get: waive this finding until 2026-12-28 owner: platform-team reason: repository migration The rule stays. The exception expires. That's a very different thing from changing the architecture policy permanently. Baselines solve a different problem I also don't want to confuse waivers with baselines. A baseline answers: This violation already existed. A waiver answers: This active violation is intentionally allowed for a limited time. And a policy change answers: We changed what the project considers acceptable. Those are three different states. That separation might sound overly precise. In practice, it makes the tool much more useful. A real codebase can have old architectural debt. It can have temporary migration exceptions. And it can deliberately evolve its architecture. If all three become "ignore this finding", you lose important information. What the agent should actually do This doesn't mean agents have to stop making architectural changes. Quite the opposite. I want them to do more. A useful agent flow could look like this: requirement ↓ agent plans change ↓ implementation ↓ deterministic verification ↓ failure ↓ agent explains why ↓ policy proposal / waiver proposal ↓ approval ↓ verification against new state The agent can drive most of that workflow. It can inspect the repository. It can understand the requirement. It can implement the refactor. It can identify the policy that stopped the change. It can prepare the evidence for a policy proposal. It can even tell me that the current architecture appears to be the problem. What it shouldn't get is an invisible path from: my change failed to: therefore my own change is now allowed This also changes what "autonomous" means I've started thinking that autonomy isn't really one switch. An agent can have permission to: - read the repository; - modify source files; - run tests; - inspect dependencies; - propose policy changes; without having permission to: - approve those policy changes; - remove its own enforcement; - extend a temporary waiver indefinitely; - change the authority that verifies it. That gives you a more useful permission model than simply: autonomous = yes/no Different operations can have different authorities. And that matters more once the agent is running for a long time or can delegate work to other agents. The verifier should not care who made the code change There's another distinction here that I find useful. Verification should primarily answer: Does the current change fit the current approved policy? It doesn't need to decide whether the author was: human Claude Codex Cursor another agent automation That's a separate concern. The verifier checks the state. The policy layer defines the constraint. The authority layer controls who can change the constraint. Keeping those pieces separate makes the system much easier to reason about. This is what I started building into Guard This distinction ended up affecting the design of Codapult Guard quite a bit. Guard already had the idea of: facts ↓ policy ↓ verification But that isn't enough when policy itself can change. So I started treating policy changes as first-class operations. Guard can discover project facts and prepare proposals. A project can explicitly approve them. For protected projects, policy approval can requ
Comments
No comments yet. Start the discussion.