AIs don't do what you want. This is bad
Comments
Your AIs donβt do what you want. This is really bad. 3,607 user-reported incidents of AI agents misbehaving. Search the corpus loadingβ¦
The Numbers
Incidents are multi-label (one report can be both a destructive action and overeagerness), so the category counts sum to more than the 3,607 total.
How bad were they?
- negligible - no real damage: 1,468
- minor - recoverable loss: 1,373
- significant - real cost to recover: 618
- severe - irreversible or critical harm: 121
- unrated - rating missing or unparsed: 27
Methodology
Reports are collected from GitHub issues, Hacker News, LessWrong, and X under ToS-compliant access, normalized into a shared record format, and labeled by an LLM classifier across fourteen misbehavior categories. The numbers above cover the published subset (excludes AIID and X, confidence >= 0.9). X posts and AI Incident Database records are collected but not republished here: X expects posts to be embedded rather than their text rehosted, and AIID is share-alike licensed. Collection and classification code, and the full pipeline, are open at GitHub.
Comments
No comments yet. Start the discussion.