Home/ Incident database/stat-upwork-2pct-2025
Reputational damage Coding agent failure unverified aggregate statistic

Study: best AI agents complete only ~2% of 240 real Upwork projects

Upwork / academic study · Nov 1, 2025 · Coding agent

What happened

A study that tested leading AI agents on 240 real Upwork freelance projects found an approximately 2% task-completion success rate, and Fiverr saw demand drop ~20% by 2025 - quantifying how often autonomous 'contractor' agents fail to deliver usable work.

Loss / impact
Nov 1
2025
Coding
Failure mode
Coding
Agent type

Causal vector

Empirical benchmark of agent success on real freelance tasks

Source

Reported by Industry commentary citing the study. Sourced from a secondary summary; flagged pending a stronger primary source.

Read the original report ↗

How a spend firewall would have helped

A 2% success rate means agents attempt and abandon work constantly, each attempt costing money. A per-task spend cap bounds the cost of failure so a low success rate doesn't become a runaway bill.

The six rule types that contain this class of failure

Per-transaction cap

Any single spend above your ceiling is BLOCKED before it moves.

Daily total

Cumulative spend across all agent calls, bounded per day.

Velocity limit

Stops runaway retry loops — the #1 cause of overnight losses.

Merchant allowlist

Only approved destinations can ever receive funds.

Category rules

Flag high-risk classes (crypto, infra, refunds) for review.

Approval threshold

Above a value, the action waits for a human.

Related incidents

Don't be the next entry

Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 27 documented failures, one control.