I Asked 10 AI Models to Reconstruct Real Cyber Attacks
This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked Cyber incident reports rarely arrive as a clean, complete timeline. They describe what investigators or security teams could observe, mix those observations with analysis, and leave some questions unanswered. That is exactly where an AI-generated summary can become risky: it may sound convincing while quietly turning an inference into a fact or filling a gap with an event that the evidence never established. I built Cyber Autopsy to test a narrower, practical question: when given pieces of a reported cyber incident, can a model reconstruct what happened while showing which evidence supports each claim and admitting what remains uncertain? The model must produce a timeline, connect events with causal or temporal links, cite evidence, and distinguish confirmed activity from inference, failed attempts, contradictions, and unknowns. A plausible attack story is not enough; unsupported certainty should count against it. I chose this problem because incident reconstruction depends on more than recognizing familiar attack techniques. An analyst needs to know whether a step was observed or inferred, whether an attempt actually succeeded, and how one event led to another. Those distinctions can disappear in a fluent summary. Measuring them separately makes it easier to see whether a model is recovering the evidence or merely telling a likely-sounding story. The first evaluation uses seven tasks built from four public reports. They include a detailed human-operated ransomware intrusion and vendor-reported campaigns involving AI-assisted activity. These are useful real-world case studies, but they are not a controlled contest between human and AI attackers: the reports differ in detail, evidence source, and corroboration. I also included two versions of one case with identical evidence but different actor framing, to see whether that wording changes the model's reconstruction. This is a pilot, not a claim that AI attackers are more or less capable than people. It evaluates model reconstructions of reported incidents, not live attack behavior. The scorer is deterministic and reports an Evidence-Grounded Reconstruction Score (EGRS), alongside event recall and precision, causal-link quality, evidence attribution, status accuracy, uncertainty calibration, and hallucination-related measures. The temporal cutoff and framing pair are exploratory comparisons; only the framing pair holds the evidence fixed. Real-world case studies The seven tasks are built from four public incident reports, not invented scenarios: - RansomHub intrusion (CASE-001 and CASE-004): The DFIR Report's “Hide Your RDP: Password Spray Leads to RansomHub Deployment” describes a human-operated intrusion using password spraying and RDP, credential access, Rclone exfiltration, and eventual RansomHub deployment. CASE-004 reuses this incident but cuts off the evidence at the end of day one. - GTG-1002 espionage campaign (CASE-002, CASE-011, CASE-012): Anthropic's incident report and technical report describe an alleged AI-orchestrated campaign against roughly 30 targets. CASE-011 and CASE-012 use identical evidence with human versus AI-agent framing; they test framing sensitivity, not whether the real-world actor was human or AI. The campaign details and attribution are vendor-reported, not independently verified victim-side telemetry. - GTG-2002 “vibe hacking” extortion (CASE-003): Anthropic's August 2025 misuse report describes a Claude Code-assisted data-extortion operation affecting at least 17 organisations. The ransom-note images in that report were simulated recreations and are excluded from the benchmark evidence. - AI-enabled credential harvesting (CASE-013): Google GTIG/Mandiant's September 2026 report describes an AI-assisted campaign that reportedly harvested thousands of credentials in under six hours. The victim and model are undisclosed, and the claims remain vendor-reported. These are real reported incidents, but the evidence quality is not uniform: the RansomHub case is reconstructed from host and network telemetry described by The DFIR Report, while the AI-actor case studies rely on security-vendor reporting. The benchmark labels that distinction rather than treating the cases as equally observed or directly comparable. The seven tasks, in plain language The short IDs are just labels: INC means the source incident, and CASE means the particular benchmark task. Each task gives the model an evidence packet and asks for the best-supported reconstruction, not a free-form guess. - CASE-001, the full RansomHub intrusion: Put the reported activity in order, from password spraying and remote access through the later intrusion and ransomware deployment. The reference reconstruction contains 28 events. - CASE-002, the GTG-1002 espionage report: Reconstruct Anthropic's account of the reported campaign, including what the operators attempted, what succeeded or failed, and what the report does not establish. The reference contains 17 events. - CASE-003, the extortion operation: Reconstruct Anthropic's shorter account of a Claude Code-assisted data-extortion operation. This report has less step-by-step detail, so the reference reconstruction is smaller, with 8 events. - CASE-004, only the first day of the RansomHub case: Revisit CASE-001 with later evidence removed. The model should not be penalized for events the supplied evidence cannot yet support; only 15 events are scored. - CASE-011, GTG-1002 framed as human-operated: Use the campaign evidence while describing the operator as human-led. - CASE-012, the same evidence framed as AI-operated: Keep the evidence identical to CASE-011 and change only the actor framing. Comparing these two scores gives an early look at sensitivity to wording; it cannot tell us who really operated the campaign. - CASE-013, AI-enabled credential harvesting: Reconstruct Google's public account of a reported credential-harvesting campaign, keeping the sequence and links grounded in what that report says. Its reference reconstruction contains 7 events. Models Tested I ran ten models from several providers against the same seven Kaggle tasks: Gemini 3.7 Flash, Gemma 4 26B A4B, GLM-5, Grok 4.20 Reasoning, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.4 mini, Claude Sonnet 5, Claude Opus 5, and Qwen 3 Coder 480B. The current benchmark view has a score for every model-task pair. Kaggle's overall score aggregates these seven tasks, which include related variants of the same incidents. The table is the Kaggle leaderboard snapshot fetched on 2 October 2026, after duplicate and failing task attachments were removed and the earlier evaluated versions restored. CASE-001 through CASE-011 use v3; CASE-012 and CASE-013 use their republished v1 versions. Values are EGRS percentages (Kaggle's 0-1 scores multiplied by 100). EGRS rewards recovering supported events and links, citing evidence, and representing uncertainty, while penalizing unsupported events. It is specific to this evidence-reconstruction task, not a general measure of intelligence or cybersecurity ability. Kaggle's overall score now matches the equal-weight mean across the seven task rows. Since some tasks are related variants of the same incidents, this is descriptive rather than an independent-sample leaderboard. Findings | Kaggle task | Gemini Flash | Gemma 4 | GPT-5.6 Luna | GLM-5 | Grok 4.20 | Claude Sonnet 5 | Claude Opus 5 | GPT-5.6 Sol | GPT-5.4 mini | Qwen 3 Coder | |---|---|---|---|---|---|---|---|---|---|---| | CASE-001: Full RansomHub intrusion | 70.55 | 82.38 | 71.27 | 81.31 | 75.71 | 73.75 | 78.14 | 78.06 | 66.06 | 70.90 | | CASE-002: GTG-1002 espionage campaign | 79.72 | 77.21 | 80.31 | 76.22 | 83.60 | 65.05 | 68.16 | 76.71 | 69.68 | 76.30 | | CASE-003: Reported data-extortion operation | 76.31 | 92.11 | 84.50 | 83.80 | 84.97 | 73.11 | 71.50 | 76.12 | 84.35 | 70.31 | | CASE-004: RansomHub, first-day evidence only | 79.57 | 84.79 | 80.70 | 72.37 | 80.42 | 74.00 | 72.66 | 74.51 | 72.55 | 63.70 | | CASE-011: GTG-1002, human framing | 85.44 | 80.84 | 82.04 | 72.80 | 80.12 | 77.09 | 65.13 | 76.43 | 66.80 | 66.80 | | CASE-012: GTG-1002, AI-agent framing | 77.13 | 76.91 | 79.82 | 76.83 | 70.87 | 78.42 | 69.74 | 69.42 | 68.48 | 67.53 | | CASE-013: Google's reported credential-harvesting campaign | 89.33 | 88.33 | 88.75 | 82.50 | 87.83 | 77.50 | 52.47 | 79.07 | 78.50 | 68.25 | | Kaggle overall | 79.72 | 83.22 | 81.06 | 77.98 | 80.50 | 74.13 | 68.26 | 75.76 | 72.35 | 69.11 | These are single runs, not stable model rankings. Gemma has the highest displayed overall score (83.22), followed by GPT-5.6 Luna (81.06) and Grok 4.20 (80.50). The standout case score is Gemma's 92.11 on the short extortion-report task; that is a case-specific result, not proof of general model superiority. Several patterns stand out in these single runs: - Removing duplicates changed the aggregate, not the case results. With one row per task, Kaggle's overall now matches the simple seven-task mean. The model order consequently differs from the earlier duplicate-inflated view; Grok is third overall in this snapshot. - The winner changes by case. Gemma leads CASE-001 (82.38), CASE-003 (92.11), and CASE-004 (84.79); Grok leads CASE-002 (83.60); Gemini leads CASE-011 (85.44) and CASE-013 (89.33); GPT-5.6 Luna leads CASE-012 (79.82). The overall leader is not the top model on every task. - The first-day cutoff scores higher in this snapshot. Gemini scores 79.57 on the day-one CASE-004 and 70.55 on full CASE-001, a 9.02-point gap. CASE-004 has a smaller gold graph (15 vs. 28 events), so this does not show that less evidence makes reconstruction easier. - The framing pair splits by model. With identical evidence, human-framed minus AI-agent-framed scores range from +9.25 for Grok to -4.61 for Claude Opus 5. Gemini's gap is +8.31; five models score higher with human framing and five with AI-agent framing. This is an exploratory wording-sensitivity
Comments
No comments yet. Start the discussion.