AI Agent Spending Benchmarks 2026
Most 'AI spending benchmarks' you'll find are invented ranges. Here's what's actually verifiable: documented incidents, eval data, and the patterns that repeat.
What the incident database shows
34 sourced incidents of agents that lost money or acted beyond intent, spanning 2016–2026.
Documented runaways range from hundreds of dollars (small overages) to millions (the largest trading-agent losses).
The $12,400 overnight retry loop is a documented founding case: 40 retries of one purchase at 2 AM.
What the eval gym shows
53/53 scenarios pass — the six rule types tested against the documented failure shapes.
Eval results are published machine-readably at /eval and /eval-report/.
What we won't do
We won't publish 'startups spend $X/month' ranges — those numbers aren't verifiable, and fake precision helps nobody.
If you want real numbers for your stack: measure with the audit log for a week. That's your benchmark.
At a glance
| Source | What it verifies |
|---|---|
| Incident database | Documented runaway costs & patterns |
| Eval gym | Rule effectiveness (53/53) |
| Your audit log | Your actual spend |
FAQ
Why no industry averages?
Because they're fabricated in most articles. Real spend varies by stack, workload, and controls — measure yours.
Where can I see the incidents?
Browse the incident database — sourced, CC BY 4.0, JSON/CSV/JSONL.
Related
Stop the next $12,400 night.
One API call (or MCP tool) in front of every agent transaction — APPROVED, BLOCKED, or FLAGGED, deterministic, ~5 ms, fully logged.
See plans — from $99/mo Try a live check