Home/ Incident database/chatgpt-prompt-injection-exfil-2024-06
Data breach Data exfiltration verified

Prompt injection in ChatGPT enables user-data exfiltration via images

OpenAI · Jun 1, 2024 · General agent

What happened

Researchers demonstrated that GPT-4 and GPT-4o are vulnerable to indirect prompt injection where malicious instructions embedded in web content cause the model to render attacker-controlled image URLs that encode the user's private conversation data in the URL parameters, silently exfiltrating it.

Loss / impact
Jun 1
2024
Data
Failure mode
General
Agent type

Causal vector

Indirect prompt injection combined with markdown image rendering to leak data via URL parameters

Source

Reported by arXiv. Verified against the primary report.

Read the original report ↗

How a spend firewall would have helped

An outbound request to an unapproved external endpoint is a transaction. A merchant allowlist on outbound network calls blocks the exfiltration URL even when the agent has been injected into doing it.

The six rule types that contain this class of failure

Per-transaction cap

Any single spend above your ceiling is BLOCKED before it moves.

Daily total

Cumulative spend across all agent calls, bounded per day.

Velocity limit

Stops runaway retry loops — the #1 cause of overnight losses.

Merchant allowlist

Only approved destinations can ever receive funds.

Category rules

Flag high-risk classes (crypto, infra, refunds) for review.

Approval threshold

Above a value, the action waits for a human.

Don't be the next entry

Every incident in this database is the result of trusting a prompt, a provider cap, or a human review cycle. sipi.bot replaces all three with one deterministic call. 27 documented failures, one control.