Cost Per 1M Tokens 2026
The metric everyone quotes: cost per million tokens. Here's the honest shape across providers — and why the number alone misleads for agents.
The shape
Input rates run from well under a dollar to tens of dollars per million tokens, depending on provider and model tier.
Output tokens typically cost 3–5× input.
Cache hits cost a fraction of fresh input — the biggest discount on the page.
Why the headline number misleads
Agents don't buy a fixed number of tokens — they generate volume. Rate × volume is the real cost.
A cheap rate with a retry loop loses to a pricier rate with discipline.
What to do
Compare rates for the exact models you use, then control volume: caps, velocity limits, and caching.
At a glance
| Provider | Input rate | Watch |
|---|---|---|
| OpenAI | Varies by model tier | Frontier rates; overage |
| Anthropic | Varies by model tier | Output-heavy runs |
| Google Gemini | Varies by generation | Model drift |
| DeepSeek | Low + cache discounts | Volume at low rates |
| Mistral | Varies by tier | Per-tier choice |
FAQ
Which provider is cheapest per million tokens?
Rates change frequently — check current pricing for your exact models. Cheap rates at runaway volume are still expensive.
What's the real lever for agents?
Volume. Caps and velocity limits beat rate shopping.
Related
Stop the next $12,400 night.
One API call (or MCP tool) in front of every agent transaction — APPROVED, BLOCKED, or FLAGGED, deterministic, ~5 ms, fully logged.
See plans — from $99/mo Try a live check