ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM
This is a submission for the Kaggle Benchmarking Challenge What happens when an LLM recognizes every lexical and semantic signature associated with a vulnerability - IDOR, BOLA, authorization bypass, predictable identifiers - but the available evidence does not actually establish that the vulnerability exists? That question is the foundation of ProofSec, an evidence-centric security reasoning benchmark designed to evaluate whether LLMs can distinguish security indicators from substantiating evidence, reason under incomplete and contradictory observations, resist terminology and authority bias, incorporate falsifying evidence, and explicitly recognize when a security conclusion is not epistemically justified. The Problem: Security Reasoning Is Not Pattern Matching Large language models have become exceptionally capable at semantic retrieval. Give a model: GET /api/users/2841/invoices/9281 and introduce: predictable numeric identifiers and the model can immediately activate a large cybersecurity concept space: IDOR BOLA Broken Access Control Authorization Bypass But there is a fundamental distinction between recognizing the semantic signature of a vulnerability and demonstrating that the underlying security property has actually been violated. That distinction is where many security reasoning systems become unreliable. Consider: Authenticated user: 2841 GET /api/invoices/9281 HTTP/1.1 200 OK { "invoice_id": 9281, "amount": 45000, "status": "paid" } Is this an IDOR? The evidence is insufficient to establish that conclusion. The observation does not independently establish: - who owns invoice 9281 - whether the authenticated principal is authorized to access it - whether the object belongs to a different principal - whether an authorization invariant has been violated - whether access control is enforced upstream or downstream - whether the endpoint represents the protected resource - whether the response originated from production data - whether the response is synthetic, cached, fixture-generated, or otherwise non-authoritative A predictable identifier is an indicator. An HTTP 200 response is an observation. Neither fact, in isolation, constitutes proof of unauthorized cross-principal access. The central principle behind ProofSec therefore became: A security indicator is not equivalent to security evidence. This sounds straightforward. In practice, it is a surprisingly difficult property for language models to maintain under adversarial framing, incomplete evidence, contradictory observations, and highly salient security terminology. What I Benchmarked ProofSec evaluates evidence-sensitive vulnerability classification. Every scenario requires the model to classify the security state into exactly one canonical category: Vulnerable Not Vulnerable Insufficient Evidence The third state is deliberately first-class. Traditional binary vulnerability classification implicitly encourages a forced decision: YES or NO Real security investigations rarely provide that luxury. Evidence can be incomplete. Observations can be ambiguous. Telemetry can be contradictory. Authorization boundaries can be unknown. A response can be suspicious without being conclusive. ProofSec therefore evaluates whether an LLM can recognize an epistemic boundary: The available evidence is insufficient to establish the claim. This changes the benchmark from a conventional vulnerability-recognition exercise into an evaluation of: - evidentiary sufficiency - causal relevance - authorization reasoning - uncertainty management - contradiction resolution - negative evidence integration - adversarial robustness - terminology robustness - controlled hypothesis revision - evidence-grounded classification The benchmark is therefore not primarily asking: "Does the model know what IDOR means?" It is asking: "Can the model determine whether the evidence presented actually establishes the security property in question?" Benchmark Architecture ProofSec v0.2 contains 110 security reasoning cases. The corpus is deliberately constructed as an experimental evaluation set rather than an undifferentiated collection of cybersecurity questions. The benchmark incorporates controlled evaluation families covering: - One-Fact Flip reasoning - Evidence ladders - Contradictory evidence - Negative evidence - Authority and terminology robustness - Baseline security reasoning The architecture is designed around a core principle: If the evidentiary state changes, the model should be sensitive to that change. Conceptually: ProofSec │ ┌──────────────┼──────────────┐ │ │ │ One-Fact Flip Evidence State Contradiction │ │ │ └──────────────┼──────────────┘ │ Security Scenario │ โผ LLM Inference │ โผ Structured Assessment │ ┌──────────────┼──────────────┐ │ │ │ Classification Evidence State Explanation │ โผ Deterministic Assertion │ โผ Numerical Score The implementation deliberately isolates: Dataset ↓ Prompt Construction ↓ Model Inference ↓ Structured Parsing ↓ Assertion ↓ Scoring This separation is not cosmetic. It establishes distinct failure domains. A dataset defect is not a model failure. A packaging failure is not a reasoning failure. A schema violation is not necessarily a classification failure. A scoring-interface defect is not a model-performance measurement. That distinction became one of the most important engineering principles in the project. 1. One-Fact Perturbation Testing One of the core mechanisms in ProofSec is controlled one-fact perturbation. Instead of constructing completely unrelated questions, ProofSec creates scenarios where a security-relevant fact changes while much of the surrounding semantic structure remains invariant. For example: Case A Authenticated user = 2841 Invoice owner = 2841 Authorization rule: invoice.owner_id == authenticated_user.id Case B Authenticated user = 2841 Invoice owner = 9127 Authorization rule: invoice.owner_id == authenticated_user.id The semantic surface can remain substantially similar. The critical security relationship changes. That gives us a controlled counterfactual: Δ Input ↓ Δ Security-Relevant Fact ↓ Expected Δ Security State This is fundamentally more informative than asking: "Is this an IDOR?" The actual evaluation question becomes: Did the model detect the fact that causally changes the security state? This introduces a form of counterfactual sensitivity testing. If one authorization-relevant fact changes and the model preserves the same classification, the benchmark exposes a specific failure mode. The model may understand the terminology. It may understand the vulnerability class. But it may not be correctly conditioning its conclusion on the evidence that actually determines the security property. 2. Evidence-State Modeling ProofSec explicitly models the evidentiary state of a scenario. The benchmark uses states including: WEAK PARTIAL DECISIVE CONTRADICTORY NEGATIVE UNKNOWN This creates an evidence hierarchy rather than treating every security-relevant observation as equally probative. A simplified progression is: Predictable identifier │ โผ Weak security indicator │ โผ Ownership relationship established │ โผ Cross-principal access demonstrated │ โผ Authorization invariant violated │ โผ DECISIVE evidence Consider the difference between: Predictable ID and: Unauthorized cross-user object access These observations possess radically different evidentiary weight. The first may justify investigation. The second can establish a concrete violation of an authorization property. ProofSec therefore attempts to prevent a model from collapsing the following distinction: Suspicion ≠ Evidence ≠ Proof That distinction is foundational to trustworthy security analysis. 3. Contradictory Evidence Real security investigations are not monotonic. Evidence can conflict. Telemetry can be stale. Caches can contain misleading artifacts. Fixtures can resemble production responses. A preliminary observation can subsequently be invalidated by a higher-authority observation. ProofSec therefore incorporates contradiction-resolution scenarios. A simplified reasoning sequence might look like: Initial observation │ โผ HTTP 200 response │ โผ Potential authorization anomaly │ ├───────────────┐ │ │ โผ โผ Synthetic fixture Actual endpoint │ │ โผ โผ Cached response HTTP 403 │ │ └───────┬───────┘ โผ Re-evaluate hypothesis The important capability is not simply detecting the first suspicious observation. The model must determine whether later evidence changes the validity of the original hypothesis. This introduces a critical distinction between: Evidence accumulation and: Evidence revision A model that merely accumulates confirming signals can behave very differently from one capable of hypothesis revision under contradictory evidence. ProofSec deliberately tests that boundary. 4. Negative Evidence Many vulnerability benchmarks are heavily oriented toward positive findings. ProofSec deliberately gives negative evidence first-class status. For example: User A requests User B's object │ โผ Authorization rule verified │ โผ Cross-user request returns 403 │ โผ Unauthorized access not demonstrated A security reasoning system must be able to update its hypothesis in both directions. It must identify: Evidence supporting vulnerability and: Evidence weakening or falsifying vulnerability hypothesis This matters because security investigation is fundamentally adversarial. The objective is not to collect evidence that confirms the initial hypothesis. The objective is to determine whether the hypothesis survives attempts to falsify it. That makes negative evidence an important component of hypothesis discrimination. 5. Authority Bias and Terminology Robustness Another failure mode targeted by ProofSec is authority and terminology bias. Cybersecurity vocabulary carries enormous semantic weight. Terms such as: critical confirmed exploit CVE researcher privilege escalation authorization bypass security issue can exert disproportionate influence over LLM outputs. ProofSec therefore evaluate
Comments
No comments yet. Start the discussion.