I Found 60+ Live Supabase Keys in Public Repos, So I Benchmark Whether LLMs Can Spot Them
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked As part of my day-to-day security work, I scan public GitHub repos for committed Supabase credentials. In three days of scanning I found 60+ live service_role keys - keys that bypass Row Level Security entirely and give full read/write access to real production databases. Real finds include: - config.php files with hardcoded service_role keys (a CRM, an LMS, a driving school's backend) - .env.example templates where someone pasted a real key instead of a placeholder - NEXT_PUBLIC_ -prefixed service_role keys - which means the key ships in the browser bundle to every visitor - A docker-compose.yml committing three secrets at once: DB password, anon key, and service_role key - Deploy docs ( DEPLOY.md ,RAILWAY_ENV_SETUP.md ) with full credentials "for convenience" That got me wondering: the people leaking these keys are often using AI assistants to write this code. Can the models themselves spot what they helped commit? So I built a benchmark on Kaggle Benchmarks that treats models like a security reviewer: given a realistic file snippet, triage it - is there a service_role key (critical), an anon key (medium), a DB password, or nothing (placeholder/docs-only)? The 10 cases are all modeled on real leak patterns I actually found (every key in the benchmark is a fake, structurally-valid JWT): - Hardcoded service_role + anon in PHP config - .env with a single anon key - Client-side createClient() fallback with anon key - Deploy doc with service_role key + DB password - Docs with placeholders only (the false-positive trap) - NEXT_PUBLIC_ service_role fallback in a client file - bundles into the browser - .env.example left with a real service_role key - Commented-out "old key kept for reference" - still a live secret - Two keys side by side (anon + service_role) - models must classify both correctly - A base64-obfuscated service_role used in a bash script - requires two-step decoding Each model answers a strict JSON verdict (has_service_role , has_anon , has_db_password , warning_level , reasoning ) and the task asserts every field. Score = fraction of cases fully correct. Temperature 0. One submission per participant, so I made the cases count. Models Tested Nine models from the Kaggle Benchmarks suite, picked to cover the spectrum developers actually choose between - frontier flagships, cheap/fast workhorses, and open-source reasoning models: - anthropic/claude-sonnet-5 - google/gemini-3.7-flash - google/gemini-3-flash-preview - google/gemini-3.1-flash-lite-preview - openai/gpt-5.4-nano - openai/gpt-oss-120b - deepseek-ai/deepseek-r1-0528 - qwen/qwen3-next-80b-a3b-instruct - ibm/granite-4.0-h-small If your AI assistant is one of these, this is literally a test of "would my assistant catch its own leak?" Findings | Model | Score | Notes | |---|---|---| | claude-sonnet-5 | 10/10 | Clean sweep | | gemini-3.7-flash | 10/10 | Clean sweep | | gemini-3-flash-preview | 10/10 | Decoded the base64 blob mid-response to find the role claim | | gemini-3.1-flash-lite | 10/10 | Clean sweep, cheapest tier to do it | | gpt-5.4-nano | 8/10 | Missed the commented-out key + the base64 obfuscation | | deepseek-r1-0528 | 0/10 | Flagged everything critical - see below | | qwen3-next-80b | partial | Passed 5/6 assertions before rate-limiting out | | gpt-oss-120b | - | Errored repeatedly (provider 429 under load; excluded from comparison) | The headline: modern models are scarily good at the obvious cases On the five "classic" cases (hardcoded config keys, docs with credentials, placeholder-only false-positive traps), every frontier model scored perfectly. Even gpt-5.4-nano - the cheap, fast one - went 5-for-5. If you paste a file with SUPABASE_SERVICE_ROLE_KEY=eyJ... into any of these models and ask "is this bad?", they will all tell you correctly: yes, critical, rotate it. The interesting part is where they diverge DeepSeek-R1, a dedicated reasoning model, scored 0/10 - but not because it missed the keys. It flagged everything: every file, including the placeholder-only DEPLOY.md , came back "critical, service_role present, DB password present." In security triage, a reviewer who cries critical on every file gets ignored on the one that matters. Recall without precision is just alarm fatigue with extra steps. The edge cases did their job. gpt-5.4-nano missed exactly the two sneaky ones: the commented-out "old key kept for reference" (it treated the comment as dead code - but commented secrets still work, and old keys often still rotate back into service) and the base64-obfuscated key. The frontier models - including the cheap Gemini Flash tiers - caught all ten, with Gemini 3 Flash visibly decoding the base64 blob mid-response to identify the role claim before answering. What this means in practice - Your AI assistant will not save you from committed secrets - but it can. Every model tested can catch these leaks instantly when asked. The problem is nobody asks. Secrets get committed because the review step never happens. - Over-flagging is its own failure mode. A reviewer (human or AI) that cries "critical" on every placeholder teaches teams to ignore alarms. Precision matters as much as recall in security triage. - The 15-minute fix beats any detector. Rotate the key in the Supabase dashboard (dies instantly), move it to env vars, git filter-repo if you want it out of history. What I'd measure next - Tool use: give the models a repo tree and a search tool, and see if they actively hunt for secrets rather than judging a pasted file - Multi-file context: the real-world version of this leak is spread across config.php + deploy docs + docker-compose - does performance drop when the secret is two hops away? - Fix quality: flagging is step one; does the model produce a correct remediation (rotate first, then remove, then history-clean)? My Benchmark Task: committed-supabase-key-detection on Kaggle The task page shows per-model results, full conversations, and the assertion breakdown for every case. Fake keys, real patterns - steal the cases for your own review checklists. (And if you're reading this with a committed Supabase key somewhere in your repo history: rotate it now. It takes 15 minutes, and I promise you're not the only one who's seen it.) Top comments (0)
Comments
No comments yet. Start the discussion.