Benchmark · 2026

AI Agent Cost per Task

What does a single completed agent task actually cost? Median across production agents: $0.02 for simple classification, $0.34 across all task types, $4.50 for multi-step research. The top 1% of runaway tasks exceed $50.

By the sipi.bot engineering team · Published 2026-07-18 · Methodology

$0.34median cost per task
$0.02–$4.50range by task type
$50+top 1% runaway tasks

Cost per task by type

Task typeMedian cost90th percentileTypical tool calls
Simple classification / routing$0.02$0.050–1
Single extraction / summary$0.08$0.201–2
Tool-augmented Q&A$0.40$1.103–6
Coding task (Claude Code, Cursor)$0.85$3.205–15
Multi-step research with tool calls$4.50$12.0015–30
Autonomous procurement / workflow$6.20$28.0010–40

What drives the cost

Three factors explain 90%+ of the variance in per-task cost:

The runaway tail. Although the median task costs $0.34, the long tail is where the money actually goes. Retry-loop waste is where the money concentrates. A research task that retries 40 times costs 40x its single-run price, and does so in seconds. A pre-spend firewall's velocity rule kills this pattern on the 11th retry.

Cost-per-task over time

Median cost-per-task has fallen ~40% over the last 12 months as providers cut prices and teams moved to cheaper models for routine work. But the runaway tail has not improved — if anything, more capable agents that chain more tool calls have made the tail worse. The savings from cheaper models are real; the savings from avoiding runaway spend are larger and easier to capture.

Methodology

  1. Anonymized sipi.bot customer data — per-task cost from teams running the firewall, aggregated by task category.
  2. Public provider pricing as of Q2 2026 for the model-cost component.

"Task" is defined as a single user-visible unit of work (one classification, one summary, one research report, one procurement). Internal retries and failed attempts are included in the cost. Released under CC BY 4.0 — cite sipi.bot.

How to use this benchmark

  1. Find your task type. Compare your median cost to the benchmark for that category.
  2. Measure your tail. If your top 1% of tasks is more than 20× your median, you have a runaway-loop problem — install a velocity rule.
  3. Right-size your models. Most classification tasks don't need GPT-4o. Most research tasks do.
Cap your cost-per-task with sipi.bot →