Home/Blog
23 posts

The sipi.bot blog

How AI agents spend money, what happens when they get it wrong, and how to build the guardrail before the bill.

Database

34 incidents, $2.91 billion tracked — the open AI agent incident database is live

We launched an open, sourced database of every documented AI agent that lost money, deleted data, or took unintended actions. Here's what the first 34 records tell us — and why we built it.

2026-07-27
Runaway Loop

The anatomy of a runaway AI agent loop

A single unbounded while-loop cost an engineering team $47,000 overnight. Here's exactly how it happens — and why velocity limits are the only reliable defense.

2026-07-14
Architecture

Why prompt instructions aren't spending controls

Prompts are suggestions, not constraints. Every documented runaway-agent incident shares the same root cause: the operator trusted what they told the model instead of what the model could do.

2026-06-02
Incidents

Step Finance $27M post-mortem: what a per-transaction cap would have done

The largest documented AI-adjacent financial incident of 2026, broken down by the firewall rule that would have contained it — and why every agent touching money needs one.

2026-05-19
Best Practices

The three false beliefs that cost agent teams money

"I'll catch it in the morning." "The provider's monthly cap will stop it." "My agent would never do that." Three beliefs, three incident stories, and what replaces them.

2026-04-15
Velocity

Velocity limits: the one rule that prevents overnight disasters

Every overnight runaway loop in our database — from $4,200 Pinecone bills to $47,000 token burns — would have been stopped cold by a velocity limit. Here's how they work and how to set one.

2026-03-10
Mcp

MCP-native spend controls: why agent tools need a payment firewall

As MCP becomes the standard for agent-tool communication, every tool that touches money needs a spend gate that lives outside the model. Here's the architecture and the integration.

2026-01-20
Agentic Payments

The state of agentic payments in 2026

x402, AP2, and AgentKit made agent-to-agent payments real. The rails are live; the control layer is the constraint. Here's where the stack stands and what ships next.

2026-08-08
Cost

How much does an AI agent actually cost to run?

The honest cost model for an AI agent: inference, tools, data, and the runaway risk that the sticker price never shows. Plus the numbers to budget.

2026-08-01
Incidents

Q2 2026 agent incident report: what the database shows

A quarterly read of the AI Agent Incident Database: the failure modes that keep repeating, the dollars involved, and the rules that would have stopped each one.

2026-07-30
Security

The hidden cost of prompt injection: it's the spend, not the prompt

Prompt injection isn't just a data-safety problem — it's a money problem. The documented attacks that ended in spend, and the deterministic defense that can't be injected.

2026-07-22
Operations

How to run an agent spend audit in 30 minutes

The audit-log questions that turn agent spend into decisions: where the money goes, what got blocked that shouldn't have, and the one rule change that matters most.

2026-07-15
Best Practices

What a spend firewall won't do

The honest limits: it won't stop every bad decision, it can't reverse settlements, and it's not a compliance certification. Here's what it actually is.

2026-07-08
Onboarding

How to pick your first three firewall rules

You don't need all six rule types on day one. The three that cover the most documented failure modes — and how to set them in one session.

2026-06-28
Voice

The voice agent spend problem nobody is talking about

Voice agents bill per minute of telephony plus LLM tokens per turn. A retry storm is billable minutes for nothing. Here's the math and the control.

2026-06-15
Operations

Subscription sprawl: the quiet AI cost

Unused seats, auto-renewing tools, silent tier upgrades — and agents that can subscribe on their own. How the quietest line on the AI bill grows.

2026-05-30
Coding Agents

The coding agent cost war is really a control problem

Every coding agent — Cline, Roo, Cursor, Claude Code, Codex — bills differently, and all of them can spend beyond the editor. The winner isn't the cheapest model; it's the team that controls the spend.

2026-05-10
Operations

Flag vs block: when to let an agent ask instead of stop

Hard blocks stop bad spend but also stop legitimate work. Approval thresholds keep agents fast. Here's the decision rule for when to use which.

2026-04-20
Architecture

Metering platforms bill your customers. They don't control your agents.

Metronome, Orb, and Amberflo are for revenue — measuring usage and billing customers. Agent spend control is a different job, in the other direction. Here's the split.

2026-04-05
App Builders

The credit model of AI app builders, and the bill it hides

Lovable, Bolt, and friends bill by credits for building. The real bill starts when the apps you built start running agents. How to budget the runtime, not just the build.

2026-03-15
Architecture

Bedrock, Vertex, and AI Foundry are not spend firewalls

The big three cloud AI platforms govern model access, not agent spend. Their budgets stop at the cloud boundary. Here's the money-layer gap.

2026-02-28
Cost

The real GPT vs Claude API cost comparison

Rate cards are the smallest part of the comparison. For agent workloads, volume, reasoning depth, and context dominate. Here's the comparison that matters.

2026-02-10
Cost

Why your AI bill keeps growing (and it's not the model)

The bill grows because of what agents do, not what models cost: tool fan-out, retry loops, context bloat, and ungoverned vendors. The five growth drivers and the five levers.

2026-01-30

Every post is backed by real data

The blog draws on the AI Agent Incident Database — 34 sourced records of agents that lost money, deleted data, or took unintended actions. No hypotheticals, no fabricated benchmarks.