Home/ Incident database/o3-arc-agi-cost-2024-12
Financial loss API cost spike verified

OpenAI o3 ARC-AGI benchmark run estimated to cost ~$346,000+ in compute

OpenAI · Dec 20, 2024 · Research agent

What happened

OpenAI's o3 reasoning model achieved a breakthrough score on the ARC-AGI benchmark, but outside analysis estimated the high-efficiency runs cost roughly $41k-$2.5M per hour and that the single semi-private benchmark run may have cost ~$346,064 in compute - a widely cited data point on the extreme per-task cost of frontier reasoning agents.

$346K
Loss / impact
Dec 20
2024
API
Failure mode
Research
Agent type

Causal vector

Massive per-task inference compute for a frontier reasoning agent with no cost ceiling

Source

Reported by Hacker News (analysis of ARC-AGI run). Verified against the primary report.

Read the original report ↗

How a spend firewall would have helped

A per-transaction cap modeled on expected task cost would FLAG any single task whose compute projection exceeds the ceiling, forcing a human to authorize frontier-scale spend before it happens.

The six rule types that contain this class of failure

Per-transaction cap

Any single spend above your ceiling is BLOCKED before it moves.

Daily total

Cumulative spend across all agent calls, bounded per day.

Velocity limit

Stops runaway retry loops — the #1 cause of overnight losses.

Merchant allowlist

Only approved destinations can ever receive funds.

Category rules

Flag high-risk classes (crypto, infra, refunds) for review.

Approval threshold

Above a value, the action waits for a human.

Don't be the next entry

Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 27 documented failures, one control.