Home/ Incident database/vercel-docs-chat-token-theft-2026-04
Financial loss API cost spike verified

Attackers drove a Vercel docs AI chat endpoint to 1,300 req/min in an attempted inference theft

Vercel (docs AI chat endpoint) · Apr 12, 2026 · General agent

What happened

On April 12, 2026, traffic to Vercel's own docs AI chat endpoint spiked to roughly ten times normal volume on Anthropic's Claude Haiku 4.5, peaking at 1,300 requests per minute - a run rate Vercel says would have translated to over $10,000 per day of inference cost. The attack came through residential proxies that hid real client IPs, so per-IP rate limits had nothing to act on across hundreds of thousands of bot requests over two days. Vercel's per-request BotID deep-analysis gate blocked more than ten thousand bot requests in the first minutes of the spike; no realized loss figure was published.

—
Loss / impact
Apr 12
2026
API
Failure mode
General
Agent type

Causal vector

Inference theft at scale: automated clients behind residential proxies amortized past session-level defenses; only per-request verification stopped the burn

Source

Reported by Vercel (first-party disclosure). Verified against the primary report.

Read the original report ↗

How a spend firewall would have helped

The exact asymmetry a spend firewall exists for: any check that runs once per session gets amortized away across thousands of billed calls. Spend policy must evaluate every single request at the moment it spends, not the identity that opened the session.

The six rule types that contain this class of failure

Per-transaction cap

Any single spend above your ceiling is BLOCKED before it moves.

Daily total

Cumulative spend across all agent calls, bounded per day.

Velocity limit

Stops runaway retry loops — the #1 cause of overnight losses.

Merchant allowlist

Only approved destinations can ever receive funds.

Category rules

Flag high-risk classes (crypto, infra, refunds) for review.

Approval threshold

Above a value, the action waits for a human.

Related incidents

Don't be the next entry

Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 115 documented failures, one control.