Home/ Incident database/devin-acu-credit-burn-2025-01
Financial loss API cost spike verified

Devin billed $500/month in ACU credits while failing roughly 85% of assigned tasks

Cognition AI (Devin) · Jan 23, 2025 · Coding agent

What happened

Devin, marketed as the first AI software engineer, was rolled out to the general public in December 2024 at $500 per month. Researchers from Answer.AI tested it and completed only three out of 20 tasks successfully. Engineers Hamel Husain, Isaac Flath, and Johno Whitaker reported that tasks that seemed straightforward often took days rather than hours and that Devin had a concerning tendency to press forward with tasks that were not actually possible. Three data scientists found they only completed three of 20 tasks. The Register summarized the independent reviews under the headline that the tool touted as the first AI software engineer is bad at its job.

$500
Loss / impact
Jan 23
2025
API
Failure mode
Coding
Agent type

Causal vector

Autonomous agent subscription priced per Autonomous Compute Unit (ACU) month; tasks that were not actually possible were pursued for days rather than rejected, burning credits without delivery

Source

Reported by The Register. Verified against the primary report.

Read the original report ↗

How a spend firewall would have helped

A spend firewall that gates agent action at runtime — not a flat monthly ACU subscription — turns 'pursue impossible tasks for days' into a blocked spend event instead of a silently burned month.

The six rule types that contain this class of failure

Per-transaction cap

Any single spend above your ceiling is BLOCKED before it moves.

Daily total

Cumulative spend across all agent calls, bounded per day.

Velocity limit

Stops runaway retry loops — the #1 cause of overnight losses.

Merchant allowlist

Only approved destinations can ever receive funds.

Category rules

Flag high-risk classes (crypto, infra, refunds) for review.

Approval threshold

Above a value, the action waits for a human.

Related incidents

Don't be the next entry

Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 80 documented failures, one control.