Attackers drove a Vercel docs AI chat endpoint to 1,300 req/min in an attempted inference theft
Vercel (docs AI chat endpoint) · Apr 12, 2026 · General agent
What happened
On April 12, 2026, traffic to Vercel's own docs AI chat endpoint spiked to roughly ten times normal volume on Anthropic's Claude Haiku 4.5, peaking at 1,300 requests per minute - a run rate Vercel says would have translated to over $10,000 per day of inference cost. The attack came through residential proxies that hid real client IPs, so per-IP rate limits had nothing to act on across hundreds of thousands of bot requests over two days. Vercel's per-request BotID deep-analysis gate blocked more than ten thousand bot requests in the first minutes of the spike; no realized loss figure was published.
Causal vector
Inference theft at scale: automated clients behind residential proxies amortized past session-level defenses; only per-request verification stopped the burn
Source
Reported by Vercel (first-party disclosure). Verified against the primary report.
The exact asymmetry a spend firewall exists for: any check that runs once per session gets amortized away across thousands of billed calls. Spend policy must evaluate every single request at the moment it spends, not the identity that opened the session.
The six rule types that contain this class of failure
Per-transaction cap
Any single spend above your ceiling is BLOCKED before it moves.
Daily total
Cumulative spend across all agent calls, bounded per day.
Velocity limit
Stops runaway retry loops — the #1 cause of overnight losses.
Merchant allowlist
Only approved destinations can ever receive funds.
Category rules
Flag high-risk classes (crypto, infra, refunds) for review.
Approval threshold
Above a value, the action waits for a human.
Related incidents
Codex Windows app exhausted a weekly usage limit in about 62 hours across many parallel tasks
lty418905-blip (OpenAI Codex user) · Aug 5, 2026
Codex context auto-compaction loop rereads files, loses progress, and burns paid credits
edisonjoao1 (OpenAI Codex user) · Jul 24, 2026
OpenAI confirmed Codex background tasks were silently burning users' usage limits faster than normal
OpenAI (Codex) · Jun 30, 2026
Don't be the next entry
Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 115 documented failures, one control.