OpenAI o3 ARC-AGI benchmark run estimated to cost ~$346,000+ in compute
OpenAI · Dec 20, 2024 · Research agent
What happened
OpenAI's o3 reasoning model achieved a breakthrough score on the ARC-AGI benchmark, but outside analysis estimated the high-efficiency runs cost roughly $41k-$2.5M per hour and that the single semi-private benchmark run may have cost ~$346,064 in compute - a widely cited data point on the extreme per-task cost of frontier reasoning agents.
Causal vector
Massive per-task inference compute for a frontier reasoning agent with no cost ceiling
Source
Reported by Hacker News (analysis of ARC-AGI run). Verified against the primary report.
A per-transaction cap modeled on expected task cost would FLAG any single task whose compute projection exceeds the ceiling, forcing a human to authorize frontier-scale spend before it happens.
The six rule types that contain this class of failure
Per-transaction cap
Any single spend above your ceiling is BLOCKED before it moves.
Daily total
Cumulative spend across all agent calls, bounded per day.
Velocity limit
Stops runaway retry loops — the #1 cause of overnight losses.
Merchant allowlist
Only approved destinations can ever receive funds.
Category rules
Flag high-risk classes (crypto, infra, refunds) for review.
Approval threshold
Above a value, the action waits for a human.
Don't be the next entry
Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 27 documented failures, one control.