I Built an AI Agent That Fixes Sonar and Snyk Findings.
title: "I Built an AI Agent That Fixes Sonar and Snyk Findings" published: false description: "AutoRemediate AI: a Java/Spring Boot agent where the LLM only proposes tiny patches and deterministic tools decide if they are safe." tags: java, ai, springboot, devops github project url : https://github.com/nitesh401/AutoRemediate-AI Every team has the same backlog: unused imports, unused variables, and libraries with known vulnerabilities. Boring to fix, easy to postpone, risky to fix carelessly. So I built AutoRemediate AI, an agent that fixes the safe ones for you. The twist is that it does not trust the AI to say a fix is safe. The AI proposes. Deterministic tools verify. A human approves. Why "just ask the LLM to fix everything" is a bad idea An LLM can be confidently wrong. It can break behavior, invent a library version that does not exist, or follow instructions hidden in a code comment. A passing sentence like "this change is safe" from a model is an opinion, not proof. Proof is a green build, passing tests, and a clean re-scan. So I built the agent around that idea: the model only does the thinking part, and plain code does everything that must be reliable. What it does Give it a Java/Maven repository. It will: - Clone it into an isolated workspace and create an ai/remediation/ branch (your main branch is never touched). - Run a baseline: mvn clean verify , then read the test reports. - Scan with Sonar (code quality) and Snyk (vulnerable dependencies). - Classify each finding and pick only high-confidence ones. - Ask Claude for a minimal patch, validate it, and have a second call review it. - Rebuild, re-run tests, re-scan, and compare before vs after. - Commit if the safety gate passes. Otherwise roll back and retry (max 3 attempts). - Write a Markdown/JSON report and, optionally, open a draft pull request. clone -> baseline -> scan -> classify -> for each safe finding: LLM plan -> validated patch -> review -> build -> tests -> rescan -> SAFETY GATE PASS: commit FAIL: reset, diagnose, retry / revert -> report -> optional draft PR (never auto-merged) The safety gate (the part I care about most) A fix is accepted only if all of this is true, checked by normal code with no LLM involved: - Baseline build was green, and the build is green after the fix - All tests pass, and the test count did not drop (no "fixing" by deleting tests) - The target finding is actually gone - No new Sonar issue and no new Snyk vulnerability - Every scanner that worked before still works afterwards (a failed rescan means "cannot prove", so reject) Sonar results are compared as fingerprints (rule | file | normalized message ), so deleting a line does not make every issue below it look new. Snyk results are compared by vulnerability ID plus package, ignoring the version the fix changes. Keeping the AI on a short leash - Tiny edits, not rewrites. The model returns JSON search/replace edits. The text to replace must match the file exactly once, otherwise nothing is written. - Strict file and size rules. Sonar fixes may touch only .java files, Snyk fixes onlypom.xml , with a cap on files and changed lines. - Dependency upgrades are done by code. A PomVersionUpdater changes one version number. It refuses managed versions and shared properties that would upgrade unrelated libraries. The model can only confirm risk, and the version must equal Snyk's recommendation, so it cannot invent one. - Fail closed. If the review step errors or is unclear, the patch is rejected. - Only safe categories are automated. Security findings, sensitive paths such as auth /crypto , transitive dependencies and major upgrades are reported for humans. Security, because it runs untrusted code The agent builds and tests other people's repositories, so: - One CommandRunner starts every process: allow-list ofmvn ,mvnw ,git ,snyk , no shell, timeouts, bounded output. - Child processes get a scrubbed environment, so build and test code never sees your API keys. - Repository content is treated as untrusted data (prompt injection defense) and templates substitute values in a single pass. - Pushes to main ,master ,develop andrelease are refused, and there is no merge capability at all. - Secrets live in environment variables and are redacted from logs and reports. - Sonar scans run against a scratch project, so your real project history is not overwritten. The job state machine Instead of an open-ended "agent loop", every job moves through explicit states, and illegal transitions throw: CREATED -> CLONING -> BASELINE -> SCANNING -> CLASSIFYING -> REMEDIATING -> BUILDING -> TESTING -> RESCANNING -> VERIFYING PASS -> next finding or REPORTING -> COMPLETED FAIL -> RETRY -> ... -> REVERTED If the untouched repo already fails its own tests, the job stops with BASELINE_FAILURE . The agent will not blame itself for problems that existed before it arrived. The stack Java 21, Spring Boot, Maven, Git, Sonar Web API, Snyk CLI, Anthropic Messages API (model name is configuration, not code), JUnit 5, Docker. The LLM sits behind a small LlmClient interface, so swapping providers is a configuration-level change. What I tested, and what I did not Honest status: the deterministic core (patching, pom editing, the safety gate, classification, Git isolation, command allow-list, parsers) is covered by unit tests, and I ran a behavior harness over those classes. The full run against a live Sonar server, Snyk and the Claude API is the next milestone, and I list the known limits in the README: single-module focus, full-build verification instead of targeted tests, in-memory job store, GitHub-only PRs. What I learned - Make the world measurable first. Baselines and before/after comparisons are what turn an AI suggestion into something you can trust. - Constrain the output, then validate it. A strict JSON contract plus code-level checks beats a clever prompt. - Use the LLM only where thinking is needed. Git, Maven, scanning and comparing stay boring and deterministic, which is also cheaper. - Design for rollback. Checkpoint, reset, retry a few times, stop. - Keep a human in the loop. Draft PRs, no auto-merge. Try it / code Repository: https://github.com/nitesh401/AutoRemediate-AI link If you are building agents that touch real code, I would love to hear how you decide what is safe to automate. Comment below. Top comments (0)
Comments
No comments yet. Start the discussion.