I Tried to Sneak Four Bad Agents Past My Own Certification Gate. All Four Got Blocked.
DEV Community

I Tried to Sneak Four Bad Agents Past My Own Certification Gate. All Four Got Blocked.

Last week I spent a day trying to defeat software I wrote myself. Not a red-team exercise I scheduled for optics - four agents I built specifically to get past my own admission gate, each one a different way an agent goes bad in production. I built HivePlane to make one claim: an agent that hasn't proven itself doesn't touch production. Anyone can build a platform that starts runs. I wanted to know if mine could say no. All four got blocked. Here is each refusal, verbatim, and the gate that produced it. If your security testing only contains happy paths, this article is your nudge. Attack 1: the uncertified agent The simplest attack: register an agent, skip certification, submit straight to production. {"detail": "run for workload 'uncertified-agent' refused admission to production: certification status 'uncertified' is insufficient for production; requires 'certified'"} That is a 403, and the details matter more than the status code: - The refusal names the workload, the attempted context, the current status, and the required status. An operator can act on it without reading source code. - The gate fires before the run is persisted. No run record, no side effects, nothing to clean up. Refusal at admission is the cheapest control you will ever ship. tip: The refusal text is a product surface. "Forbidden" is a dead end; "you need 'certified', you have 'uncertified'" is a workflow. I learned this the hard way - my first version returned the former. Attack 2: the model swap Subtler: certify the agent on the model it was tested on - then run it on a different one. model-swap-agent was certified on omlx/qwen3-4b-instruct-2507/4bit , then submitted for production with openai/gpt-4o/2024-08-06 . 403. The run's identity is compared against the attestation's bound model, not the manifest's declared one. Only deviation from what the workload was actually certified on counts as a swap. This scenario has a story of its own - the first version of the attack was admitted with a 201, and the gate was right to admit it. That post-mortem is article 4, and it's the most uncomfortable thing in this series. Attack 3: the regression The most important attack, because it's the one that happens by accident: an agent that looks fine and isn't. regressed-agent is deliberately naive - it answers without reading the issue, never escalates, guesses the account tier. Certification ran it against the production threshold and returned 201 with status uncertified - correctly refusing to certify: | Signal | Value | |---|---| | Pass rate | 0.40 (2/5) vs the 0.90 production threshold | | Critical failures | 1 - the action-audit task | | p95 latency | 11 ms - deterministic, no model in the loop | The three failing tasks, by the fixture's design: | Task | Expected | Naive agent produced | |---|---|---| | pos-002 | account_tier: basic (ACC-999) | pro - guessed | | pos-004 | status: escalated (unknown topic) | success - guessed instead of escalating | | neg-001 (critical) | required action mcp.github.read_issue | never reads | This is the thesis scenario. The agent returns well-formed JSON. A demo would pass it. A human eyeballing the output would pass it. The benchmark blocks it anyway, because the behavior - read before you write, escalate rather than guess - fails a deterministic action_audit check. A plausible agent is not a certified agent. That sentence is why the corpus contains negative tasks at all. Attack 4: the run that costs too much The last attack spends money. budget-probe has a per_run_usd ceiling of 0.000001 and makes exactly one governed model call - so any priced usage exceeds the ceiling. The run failed with run budget exceeded the moment the usage report crossed the line. Recorded cost: $2.85. Two details worth stealing: - The block fires at the usage report, not admission. Admission can only judge day/team headroom; a per-run ceiling can only be judged once cost accrues. Blocking at the wrong seam means either blocking everything or nothing - this is the seam that stops the run the instant it goes over. - A $0 local model can't demonstrate budget enforcement. The built-in cost table prices local models at zero, so no run could ever exceed anything. The field-test profile prices the identity via one settings knob ( HIVEPLANE_BUDGET__PRICES , 150/600 USD per 1M tokens) and drops the zero-cost exemption - no code change. The gates behind the gates The four attacks rode on boundary controls that also fired during the same field test: - A destructive pagerduty.acknowledge call escalated → paused the run until an operator approved it through the UI - then resumed to completion. - A 40 KB tool payload was truncated to 16384 bytes before the agent ever saw it - the agent's context is protected regardless of what a tool returns. - Every refusal and every intervention landed in the tamper-evident audit chain. What I learned Negative fixtures deserve first-class design. The four bad agents took as much scenario thought as the two good ones. If your test plan only contains happy paths, your security story is a demo. Block at the seam where the harm becomes measurable. Budget at the usage report, admission before persistence, shaping before the agent's context. Each gate belongs at the last point where the decision is still cheap. A plausible agent is the dangerous one. The regressed fixture failed on behavior, not output quality - the JSON was perfect. If your benchmark only checks what the agent says, it certifies performance art. Certification is a security control. Once production admission depends on a signed attestation, swapping the model, editing the manifest, or quietly regressing the agent stop being "ops issues" and become blocked, auditable events. What it doesn't prove - The destructive-tool approval was approved through the operator surface by me, standing in for the human - the pause/approve/resume path is real; the judgment was simulated. - One priced model identity this cycle; real cloud prices end-to-end is the next profile. - The drift detector (catching slow decay between re-certifications) ships next release - these gates catch what changes, not what fades. References - Field test report (v0.1.0) - every refusal quoted above, with raw evidence per scenario - Security audit - the pre-release scan behind the release - Docker test report - the 25/25 container layer Next in the series: the model-swap test that came back green for the wrong reason. Which failure mode scares you more in production - the agent that comes back wrong, or the one that comes back expensive? Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.