The voice agent spend problem nobody is talking about
Voice agents bill per minute of telephony plus LLM tokens per turn. A retry storm is billable minutes for nothing. Here's the math and the control.
Voice agents are the fastest-growing agent category — and their billing model is the easiest to overspend: per minute of telephony on top of per-token LLM cost for every turn. Both lines scale with autonomy.
The two lines
The telephony line is simple: minutes × rate. The token line is where it escapes — a long or looping conversation compounds tokens per turn, and every tool call mid-call adds another spend surface. Platforms like Vapi, Retell, and Bland bill the minutes; nobody bills the policy.
The retry storm
Failed calls get retried — and retries are billable minutes for nothing. A retry storm during an outage is the voice version of the runaway loop: same pattern, different meter. A velocity limit stops it at the source.
The control
Per-agent ceilings on telephony, category rules separating minutes from tokens, and an allowlist for mid-call tools. See the voice agent spend guide →