What is Token Caching?
Token caching lets providers reuse your context across calls and bill it at a discount — the biggest lever on the rate side of the cost equation.
How it works
Providers cache stable prefixes (system prompts, tool schemas).
Cached input tokens bill at a lower rate.
Why it matters for agents
Agents re-send large contexts every turn.
Caching turns repeated context into a discount.
The honest caveat
Cache hit rates vary by workload — measure yours.
Caching cuts the rate; it doesn't govern the volume.
FAQ
Is caching automatic?
Often, for stable prefixes — check your provider's docs.
Does caching replace caps?
No — it cuts rate; caps bound volume. Both.
Related
Stop the next $12,400 night.
One API call (or MCP tool) in front of every agent transaction — APPROVED, BLOCKED, or FLAGGED, deterministic, ~5 ms, fully logged.
See plans — from $99/mo Try a live check