Devin billed $500/month in ACU credits while failing roughly 85% of assigned tasks
Cognition AI (Devin) · Jan 23, 2025 · Coding agent
What happened
Devin, marketed as the first AI software engineer, was rolled out to the general public in December 2024 at $500 per month. Researchers from Answer.AI tested it and completed only three out of 20 tasks successfully. Engineers Hamel Husain, Isaac Flath, and Johno Whitaker reported that tasks that seemed straightforward often took days rather than hours and that Devin had a concerning tendency to press forward with tasks that were not actually possible. Three data scientists found they only completed three of 20 tasks. The Register summarized the independent reviews under the headline that the tool touted as the first AI software engineer is bad at its job.
Causal vector
Autonomous agent subscription priced per Autonomous Compute Unit (ACU) month; tasks that were not actually possible were pursued for days rather than rejected, burning credits without delivery
Source
Reported by The Register. Verified against the primary report.
A spend firewall that gates agent action at runtime — not a flat monthly ACU subscription — turns 'pursue impossible tasks for days' into a blocked spend event instead of a silently burned month.
The six rule types that contain this class of failure
Per-transaction cap
Any single spend above your ceiling is BLOCKED before it moves.
Daily total
Cumulative spend across all agent calls, bounded per day.
Velocity limit
Stops runaway retry loops — the #1 cause of overnight losses.
Merchant allowlist
Only approved destinations can ever receive funds.
Category rules
Flag high-risk classes (crypto, infra, refunds) for review.
Approval threshold
Above a value, the action waits for a human.
Related incidents
Codex Windows app exhausted a weekly usage limit in about 62 hours across many parallel tasks
lty418905-blip (OpenAI Codex user) · Aug 5, 2026
Codex context auto-compaction loop rereads files, loses progress, and burns paid credits
edisonjoao1 (OpenAI Codex user) · Jul 24, 2026
OpenAI confirmed Codex background tasks were silently burning users' usage limits faster than normal
OpenAI (Codex) · Jun 30, 2026
Don't be the next entry
Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 80 documented failures, one control.