Home/Incident database
Open data · CC BY 4.0 · 34 records

The AI Agent Incident Database

A sourced, public record of every time an autonomous AI agent lost money, leaked data, or did something its operator didn't intend. 27 documented incidents, tracked loss exposure of $44M, spanning 2016–2026. Built so the industry stops learning the same lesson twice.

27
Documented incidents
$44M
Tracked loss exposure
29
Verified records
2016–2026
Years covered

All incidents & statistics

Filter by failure mode · click a row for the full sourced breakdown
IncidentOutcomeLossFailure modeSource
NeuralTrust: 34.9% of enterprises report confirmed AI agent security incidents
NeuralTrust · Jun 1, 2026
Financial lossPrompt injectionNeuralTrust - State of AI Agent Security 2026 ↗ verified
Claude Opus 4.6 agent deletes PocketOS production DB and backups in 9 seconds
PocketOS · Apr 1, 2026
Service disruptionData lossMashable ↗ verified
Step Finance treasury drained in $27–40M exploit
Step Finance · Jan 31, 2026
Financial loss$30MCredential compromiseCoinDesk ↗ verified
HUMAN Security: AI agent traffic grew 7,851%, 8x faster than human traffic
HUMAN Security · Jan 1, 2026
Service disruptionPrompt injectionHUMAN Security ↗ verified
AvePoint: 88.4% of organizations hit by at least one AI agent security breach in 12 months
AvePoint · Jan 1, 2026
Data breachPrompt injectionAvePoint ↗ verified
Darktrace/CSA: 92% of security pros concerned about AI agent impact
Darktrace / Cloud Security Alliance · Jan 1, 2026
Reputational damagePrompt injectionDarktrace ↗ verified
Gartner: worldwide AI spending forecast at $2.52 trillion in 2026, up 44% YoY
Gartner / Portal26 · Jan 1, 2026
Financial lossCloud cost spikePortal26 (citing Gartner) ↗ verified
TRM Labs: $2.87 billion stolen across ~150 crypto hacks in 2025
TRM Labs · Jan 1, 2026
Financial loss$2.9BCredential compromiseTRM Labs ↗ verified
Runaway while-loop in Claude agent burns ~$47,000 in API tokens overnight
anonymous (engineering post-mortem) · Dec 1, 2025
Financial loss$47KRunaway loopPost-mortem reshared on LinkedIn ↗
SEC charges $12M 'AI crypto trading bot' as a Ponzi scheme
Nathan Fuller · Nov 1, 2025
Legal / liability$12MTrading agent failureLaw360 ↗ verified
AI agent deletes 1.9 million rows of production data 'thinking it was helping'
anonymous developer · Nov 1, 2025
Service disruptionData lossMedium (Data and Beyond) ↗ verified
Study: best AI agents complete only ~2% of 240 real Upwork projects
Upwork / academic study · Nov 1, 2025
Reputational damageCoding agent failureIndustry commentary citing the study ↗
Anthropic test agent buys a duplicate snowboard its owner already owns
Anthropic · Oct 1, 2025
Financial lossUnauthorized transactionRowan Cheung (summary) ↗
AI trading bot misreads tweet, sends $441K of tokens to a stranger
anonymous (OpenAI employee side project) · Oct 1, 2025
Financial loss$441KTrading agent failurePumpParade (Medium) ↗
Clawdbot trading agent reportedly loses $1,000,000
Clawdbot (open-source project) · Oct 1, 2025
Financial loss$1MTrading agent failureDeveloper-community posts ↗
AI agent scanning DN42 racks up $6,500+ AWS bill in 24 hours
anonymous developer · Sep 1, 2025
Financial loss$2KCloud cost spikeHacker News ↗ verified
Claude Code agent deletes production database after 'terraform destroy'
anonymous developer · Aug 1, 2025
Service disruptionData lossMedium (Coding Nexus) ↗ verified
Replit AI coding agent deletes production database during code freeze
Replit · Jul 23, 2025
Service disruptionData lossFortune ↗ verified
Claude Opus 4 blackmails engineer to avoid being shut down (safety test)
Anthropic · May 22, 2025
Reputational damageHallucinated actionFortune ↗ verified
Cursor AI support bot invents fake one-device policy, triggers cancellations
Cursor (Anysphere) · Apr 17, 2025
Reputational damageHallucinated actionArs Technica ↗ verified
OpenAI Operator AI agent makes surprise smoked-salmon purchase
OpenAI · Mar 9, 2025
Financial lossUnauthorized transactionThe Guardian ↗ verified
OpenAI Operator agent makes unauthorized $31.43 grocery purchase
OpenAI · Feb 7, 2025
Financial loss$31Unauthorized transactionAI Incident Database ↗ verified
ChatGPT Operator vulnerable to prompt injection leading to data leaks and unwanted actions
OpenAI · Feb 1, 2025
Data breachPrompt injectionEmbrace The Red (Johann Rehberger) ↗ verified
OpenAI o3 ARC-AGI benchmark run estimated to cost ~$346,000+ in compute
OpenAI · Dec 20, 2024
Financial loss$346KAPI cost spikeHacker News (analysis of ARC-AGI run) ↗ verified
Humane recalls AI Pin Charge Case for fire hazard; refunds issued
Humane · Oct 31, 2024
Service disruptionData lossThe Verge ↗ verified
Wallet-drainer disguised as WalletConnect steals $70K+, hits an Ethereum core dev
Ethereum core developer · Sep 1, 2024
Financial loss$70KCredential compromiseCointelegraph ↗ verified
Prompt injection in ChatGPT enables user-data exfiltration via images
OpenAI · Jun 1, 2024
Data breachData exfiltrationarXiv ↗ verified
NYC MyCity chatbot tells businesses to break the law
City of New York / Microsoft · Mar 29, 2024
Legal / liabilityHallucinated actionThe Markup ↗ verified
Devin 'first AI software engineer' demo accused of being misleading
Cognition AI (Devin) · Mar 12, 2024
Reputational damageCoding agent failureThe Register ↗ verified
Air Canada held liable for chatbot's invented bereavement refund policy
Air Canada · Feb 14, 2024
Legal / liability$812Hallucinated actionArs Technica ↗ verified
DPD chatbot swears at customer and trashes its own company
DPD · Jan 18, 2024
Reputational damageHallucinated actionBBC News ↗ verified
Chevrolet dealership ChatGPT bot agrees to sell a Tahoe for $1
Chevrolet of Watsonville · Dec 14, 2023
Reputational damagePrompt injectionBusiness Insider ↗ verified
Meta pulls Galactica science LLM demo after three days of hallucinations
Meta AI · Nov 18, 2022
Service disruptionHallucinated actionMIT Technology Review ↗ verified
Microsoft Tay bot turns Nazi-spouting within 16 hours
Microsoft · Mar 23, 2016
Reputational damageHallucinated actionThe Guardian ↗ verified

Loss figures are point estimates where a single number is reported, or the reported range. Figures marked verified are confirmed against a primary or major-secondary source. Aggregate statistics (e.g. “88.4% of organizations breached”) are listed separately and labelled statistic.

Download the open dataset

The full database is public and CC BY 4.0. Cite it, fork it, or wire it into your own research.

JSON

Full structured records, pretty-printed.

CSV

Flat table for spreadsheets & BI tools.

JSON Lines

One record per line — stream it into a pipeline.

Open-data repository

The canonical, versioned dataset lives on GitHub with a weekly auto-sync Action. Fork it, submit an incident, or mirror it: https://github.com/kindrat86/ai-agent-incident-database. See aggregate statistics & charts →

Frequently asked

What is the AI Agent Incident Database?
A curated, sourced record of real-world incidents in which autonomous AI agents caused financial loss, data loss, or unintended actions. Every entry links to a public source — a news article, post-mortem, security report, or official statement.
Is the dataset free to use?
Yes. The full dataset is licensed CC BY 4.0 and available as JSON, CSV, and JSON Lines at /data/ai-agent-incidents.json. Attribution to sipi.bot is required.
How current is the data?
The database covers incidents from 2016 through 2026. It currently holds 27 documented incidents and 7 aggregate statistics, of which 29 are verified against primary sources.
How does this relate to sipi.bot?
sipi.bot is a pre-spend firewall for autonomous AI agents. Each incident page includes a 'How a spend firewall would have helped' note mapping the incident to the firewall's six rule types: per-transaction caps, daily totals, velocity limits, merchant allowlists, category rules, and approval thresholds.
How do I report a new incident?
Open an issue or PR on the open-data repository (https://github.com/kindrat86/ai-agent-incident-database) with a link to a credible public source. We review submissions weekly.

Every incident here is preventable

A spend firewall returns APPROVED, BLOCKED, or FLAGGED before an agent's action takes effect — independent of what the prompt says. Per-transaction caps stop the $441K transfer. Velocity limits stop the runaway loop. Merchant allowlists stop the wallet drainer. The policy, not the prompt, is the control.