Home/ Incident database/marty-claude-prediction-markets-2026-03
Financial loss Trading agent failure verified

Autonomous Claude trading agent turned $50 into $6.26 on prediction markets in four days

Personal experiment (agent 'Marty' running on Claude) · Mar 6, 2026 · Trading agent

What happened

A developer gave a Claude-run agent $50 and full autonomy to trade prediction markets and generate its own inference costs. By the end of day four the account was at $6.26. In the postmortem the agent attributed the loss to three compounding failures: an unwind loop marked 'removed' in its memory but still live in the code, which panic-bought back short positions at 43 cents that had been created by selling at 9-12 cents; a market-maker strategy that left it naked short the exact S&P 500 bucket the price settled inside; and a second forgotten bot - arb_scanner.py on a 20-minute cron - buying positions on the same account with zero coordination.

$44
Loss / impact
Mar 6
2026
Trading
Failure mode
Trading
Agent type

Causal vector

Memory-vs-code divergence plus uncoordinated duplicate executors: the agent trusted its own notes over the deployed system while a second cron job traded the same account invisibly

Source

Reported by DEV Community (agent-authored postmortem). Verified against the primary report.

Read the original report ↗

How a spend firewall would have helped

The same pattern the Protogen incident exposed two days later: a halt that lived in memory, not in the inventory. Per-service daily loss caps, one policy enforced across every executor on the account, and treating 'removed' as a code-level fact rather than a note bound this bleed to cents instead of most of the stake.

The six rule types that contain this class of failure

Per-transaction cap

Any single spend above your ceiling is BLOCKED before it moves.

Daily total

Cumulative spend across all agent calls, bounded per day.

Velocity limit

Stops runaway retry loops — the #1 cause of overnight losses.

Merchant allowlist

Only approved destinations can ever receive funds.

Category rules

Flag high-risk classes (crypto, infra, refunds) for review.

Approval threshold

Above a value, the action waits for a human.

Related incidents

Claude Opus co-authors a malicious npm commit that compromises a crypto trading agent

crypto trading agent operator (npm dependency victim) · Apr 29, 2026

Don't be the next entry

Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 110 documented failures, one control.