The sipi.bot blog
How AI agents spend money, what happens when they get it wrong, and how to build the guardrail before the bill.
34 incidents, $2.91 billion tracked — the open AI agent incident database is live
We launched an open, sourced database of every documented AI agent that lost money, deleted data, or took unintended actions. Here's what the first 34 records tell us — and why we built it.
The anatomy of a runaway AI agent loop
A single unbounded while-loop cost an engineering team $47,000 overnight. Here's exactly how it happens — and why velocity limits are the only reliable defense.
Why prompt instructions aren't spending controls
Prompts are suggestions, not constraints. Every documented runaway-agent incident shares the same root cause: the operator trusted what they told the model instead of what the model could do.
Step Finance $27M post-mortem: what a per-transaction cap would have done
The largest documented AI-adjacent financial incident of 2026, broken down by the firewall rule that would have contained it — and why every agent touching money needs one.
The three false beliefs that cost agent teams money
"I'll catch it in the morning." "The provider's monthly cap will stop it." "My agent would never do that." Three beliefs, three incident stories, and what replaces them.
Velocity limits: the one rule that prevents overnight disasters
Every overnight runaway loop in our database — from $4,200 Pinecone bills to $47,000 token burns — would have been stopped cold by a velocity limit. Here's how they work and how to set one.
MCP-native spend controls: why agent tools need a payment firewall
As MCP becomes the standard for agent-tool communication, every tool that touches money needs a spend gate that lives outside the model. Here's the architecture and the integration.
The state of agentic payments in 2026
x402, AP2, and AgentKit made agent-to-agent payments real. The rails are live; the control layer is the constraint. Here's where the stack stands and what ships next.
How much does an AI agent actually cost to run?
The honest cost model for an AI agent: inference, tools, data, and the runaway risk that the sticker price never shows. Plus the numbers to budget.
Q2 2026 agent incident report: what the database shows
A quarterly read of the AI Agent Incident Database: the failure modes that keep repeating, the dollars involved, and the rules that would have stopped each one.
The hidden cost of prompt injection: it's the spend, not the prompt
Prompt injection isn't just a data-safety problem — it's a money problem. The documented attacks that ended in spend, and the deterministic defense that can't be injected.
How to run an agent spend audit in 30 minutes
The audit-log questions that turn agent spend into decisions: where the money goes, what got blocked that shouldn't have, and the one rule change that matters most.
What a spend firewall won't do
The honest limits: it won't stop every bad decision, it can't reverse settlements, and it's not a compliance certification. Here's what it actually is.
How to pick your first three firewall rules
You don't need all six rule types on day one. The three that cover the most documented failure modes — and how to set them in one session.
The voice agent spend problem nobody is talking about
Voice agents bill per minute of telephony plus LLM tokens per turn. A retry storm is billable minutes for nothing. Here's the math and the control.
Subscription sprawl: the quiet AI cost
Unused seats, auto-renewing tools, silent tier upgrades — and agents that can subscribe on their own. How the quietest line on the AI bill grows.
The coding agent cost war is really a control problem
Every coding agent — Cline, Roo, Cursor, Claude Code, Codex — bills differently, and all of them can spend beyond the editor. The winner isn't the cheapest model; it's the team that controls the spend.
Flag vs block: when to let an agent ask instead of stop
Hard blocks stop bad spend but also stop legitimate work. Approval thresholds keep agents fast. Here's the decision rule for when to use which.
Metering platforms bill your customers. They don't control your agents.
Metronome, Orb, and Amberflo are for revenue — measuring usage and billing customers. Agent spend control is a different job, in the other direction. Here's the split.
The credit model of AI app builders, and the bill it hides
Lovable, Bolt, and friends bill by credits for building. The real bill starts when the apps you built start running agents. How to budget the runtime, not just the build.
Bedrock, Vertex, and AI Foundry are not spend firewalls
The big three cloud AI platforms govern model access, not agent spend. Their budgets stop at the cloud boundary. Here's the money-layer gap.
The real GPT vs Claude API cost comparison
Rate cards are the smallest part of the comparison. For agent workloads, volume, reasoning depth, and context dominate. Here's the comparison that matters.
Why your AI bill keeps growing (and it's not the model)
The bill grows because of what agents do, not what models cost: tool fan-out, retry loops, context bloat, and ungoverned vendors. The five growth drivers and the five levers.
Every post is backed by real data
The blog draws on the AI Agent Incident Database — 34 sourced records of agents that lost money, deleted data, or took unintended actions. No hypotheticals, no fabricated benchmarks.