DEV Community

Your agent isn't reckless. It just can't see the blast radius.

I've been running Claude Code as a daily driver for about three months now. It writes Ansible I'd have taken a week to write. It reads a codebase faster than I do. It is, genuinely, very good. It also once wanted to force-push to main , and it wanted to for an extremely good reason. Sit with that for a second, because it's the whole post. The rebase was stuck. Force-pushing would have unstuck it. Every link in that chain of reasoning is sound. The agent wasn't being careless, wasn't hallucinating, wasn't "drifting" or whatever we're calling it this month. It made a locally correct decision with a non-local consequence, which is the exact category of mistake that human code review is worst at catching - because the diff looks fine. It could see the command. It could not see the crater. The thing I stopped doing For a while my answer was to read everything. Every diff, every command, eyes on the screen, hand hovering over Ctrl-C like a man watching a toddler near a staircase. This does not scale, and the reason it doesn't is embarrassing when you say it out loud: reviewing output scales with how much the agent writes. That number is going exactly one direction, and it isn't down. So I flipped it. Instead of reviewing what it produces, I started writing down what it must never do. And here's the good news that took me way too long to notice: that list is short. Not "short for a security policy" short. Short like you can fit it on a napkin. Here's mine: - A credential it read an hour ago gets inlined into a source file. - A rebase gets stuck, and the fastest route to a green terminal is git push --force origin main . - rm -rf "$BUILD_DIR/" runs on the one machine whereBUILD_DIR never got set. - A version bump gets typed straight into package-lock.json , because that's the file the version number is visibly in. - A failing test quietly grows a .skip and CI goes green. - Someone runs cat .env "just to see which variables exist." That last one is my favourite, and I'll come back to it. None of these are the agent being stupid. Every single one is a reasonable move by something that can't see two feet past the command it's about to run. Claude Code will let you say no This is the part I think a lot of people don't know exists. PreToolUse is a hook that fires before any tool call. Your script gets the whole thing on stdin: { "session_id": "abc123", "cwd": "/home/rabih/app", "hook_event_name": "PreToolUse", "tool_name": "Bash", "tool_input": { "command": "git push --force origin main" } } And you can refuse it: { "hookSpecificOutput": { "hookEventName": "PreToolUse", "permissionDecision": "deny", "permissionDecisionReason": "This force-pushes to main, a shared branch." } } Now, the bit that genuinely surprised me. That permissionDecisionReason string? The agent reads it. And acts on it. Say "blocked" and it shrugs and retries with slightly different syntax, like a cat testing a closed door. Say "change the manifest and run pnpm add " and it goes and does that, first try, no argument. Which reframes the whole thing. A denial isn't just a fence. It's the highest signal-to-noise teaching moment you will ever get, because it lands at the precise second the agent was about to be wrong. Nobody reads documentation at that moment. Everybody reads an error. So every guard I wrote has to answer two questions, not one: what's wrong, and what to do instead. Thirteen of them They live here: claude-guardrails | Guard | Blocks | |---|---| secrets-never-land-in-source | Credential-shaped literals written into source | secret-files-stay-out-of-context | Reading .env , .pem , ~/.aws/credentials into the session | secrets-are-not-staged | git add -A in a repo where .env was never gitignored | shared-branches-are-not-rewritten | git push --force to main , develop , release/ | uncommitted-work-is-not-discarded | git reset --hard , git clean -fd , git stash drop | verification-hooks-are-not-bypassed | --no-verify , HUSKY=0 , --no-gpg-sign | unexpanded-variables-in-destructive-paths | rm -rf "$DIR/" where $DIR could be empty | remote-code-is-not-piped-to-a-shell | `curl โ€ฆ \ | {% raw %}committed-migrations-are-immutable | Editing a migration that's already committed | destructive-sql-needs-a-where | Unbounded DELETE /UPDATE , ad-hoc TRUNCATE | cluster-targets-are-explicit | Destructive kubectl with no --context | tests-are-not-silenced | Introducing .skip , @Disabled , continue-on-error: true | lockfiles-are-generated-not-edited | Hand-editing package-lock.json and friends | Zero dependencies. Nothing to configure. Node reading a JSON payload and occasionally saying no. Four of them turned out more interesting than I expected when I started writing them. 1. Reading a secret is worse than writing one My first instinct was to guard the write - stop the key from landing in a file. Then I thought about it for another minute and realised I had it backwards. The write path has a code review in front of it. Someone, eventually, looks at that diff. The read path has nothing. When an agent runs cat .env to check which variables exist, it gets a completely reasonable answer to a completely reasonable question - and every value in that file is now sitting in a transcript. Transcripts get stored. Synced. Occasionally pasted into a bug report by someone being helpful. Nothing changed on disk. git diff is empty. And your credentials have left the building. So the guard blocks the read and suggests this instead: grep -o "^[A-Z_]*=" .env Same question, answered, minus the part that ruins your week. 2. "Already applied" is unknowable. "Already committed" isn't. I wanted a guard that stops you editing a migration a database has already run. Small problem: a hook has no idea what your production database has run. It's a Node script with a JSON blob. It cannot phone Postgres. But it can ask git one question: execFileSync('git', ['ls-files', '--error-unmatch', '--', pathspec], { cwd, stdio: 'ignore' }); Is this file tracked? That's it. That's the whole heuristic - and it's a good one, because once a migration is committed, something somewhere has almost certainly run it. The lovely side effect: the migration you're still drafting is untracked, so the guard is invisible while you're writing and immovable the moment you're not. The git index draws that line for free, and I didn't have to invent a single config option to get it. 3. The dangerous git add is the one that looks harmless git add .env is fine, honestly. It's visible. It's right there in the scrollback, you'd catch it. git add -A in a repo where nobody remembered to gitignore .env - that stages it silently alongside forty other files, and then the commit message says "add feature", and nobody looks, and it's on GitHub. So this guard doesn't pattern-match the command at all. It asks git what a blanket add would actually pick up: execFileSync('git', ['status', '--porcelain', '--untracked-files=all'], { cwd, encoding: 'utf8' }) Here's the part I'm quietly pleased about: gitignored files never show up in that output. Which means on a correctly configured repo, this guard is completely, permanently silent. It only ever speaks to the repos that have the problem. A guard nobody notices is a guard nobody uninstalls. That property is worth more than the check. 4. One guard just admits defeat kubectl delete pod api-7f9d . Which cluster is that? I don't know. You don't know. The agent doesn't know. The hook definitely doesn't know, because the answer lives in a config file that the payload never carries. Every other guard in this repo reads intent off the tool call. This one can't. So it does the only honest thing available: it refuses until you write --context and make the command say out loud what it's about to change. It isn't blocking a mistake. It's blocking an ambiguity - a command whose transcript won't record what it did. I think it might be the most useful one in the set, and it's the only one that works by admitting it can't see. Three rules I had to get right before I'd accept anyone else's guard If this repo works at all, most of the guards in it will eventually be written by strangers. Which changes the design problem completely. A broken guard must never block a tool call. Someone will ship a bug. If their bug takes down my git push , this whole idea dies. So every guard runs in its own try/catch , and a throw is treated as "no opinion" with a grumble on stderr. Yes - that means a crashing guard fails open. For a security tool that sounds indefensible right up until you picture the alternative: one bad merge and nobody on earth can commit until it's reverted. The plugin gets deleted, and a deleted plugin guards nothing. Fail-open keeps it installed. Installed is the entire game. Silence means allow. The dispatcher only ever emits JSON to deny. permissionDecision will happily accept "allow" , which would stomp on your own permission settings - and this plugin has no business doing that. It gets one vote. The vote is "no". Precision beats recall, and it isn't close. One false positive on correct work and the plugin is gone by lunchtime. So every guard ships its near misses as executable examples: examples: { blocked: [ { tool_name: 'Bash', tool_input: { command: 'git push --force origin main' } } ], allowed: [ { tool_name: 'Bash', tool_input: { command: 'git push --force-with-lease origin main' } }, { tool_name: 'Bash', tool_input: { command: 'git push --force origin feature/x' } } ] } --force-with-lease against --force . .env.example against .env . docs/package-lock.md against package-lock.json . That's where false positives live, so that's what you have to write down. And those examples are the test suite. npm test walks every guard and asserts both lists. That was the design decision I'm happiest with, and it took the longest to see. The obvious version of this repo has a guards/ folder and a test/ folder and contributors write both. Except they don't. Nobody writes the second folder. Ever. Foldin

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.