Study: best AI agents complete only ~2% of 240 real Upwork projects
Upwork / academic study · Nov 1, 2025 · Coding agent
What happened
A study that tested leading AI agents on 240 real Upwork freelance projects found an approximately 2% task-completion success rate, and Fiverr saw demand drop ~20% by 2025 - quantifying how often autonomous 'contractor' agents fail to deliver usable work.
Causal vector
Empirical benchmark of agent success on real freelance tasks
Source
Reported by Industry commentary citing the study. Sourced from a secondary summary; flagged pending a stronger primary source.
A 2% success rate means agents attempt and abandon work constantly, each attempt costing money. A per-task spend cap bounds the cost of failure so a low success rate doesn't become a runaway bill.
The six rule types that contain this class of failure
Per-transaction cap
Any single spend above your ceiling is BLOCKED before it moves.
Daily total
Cumulative spend across all agent calls, bounded per day.
Velocity limit
Stops runaway retry loops — the #1 cause of overnight losses.
Merchant allowlist
Only approved destinations can ever receive funds.
Category rules
Flag high-risk classes (crypto, infra, refunds) for review.
Approval threshold
Above a value, the action waits for a human.
Related incidents
Devin 'first AI software engineer' demo accused of being misleading
Cognition AI (Devin) · Mar 12, 2024
Don't be the next entry
Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 27 documented failures, one control.