How Much Does vLLM Cost?
vLLM is free software — the cost is the GPU, the power, and the engineering. Self-served inference trades token bills for a different ledger.
How vLLM is priced
No per-token fee — you pay for hardware, power, and ops.
Serving cost per million tokens depends on model size and your rig.
Time is a real cost: slower serving means longer agent loops.
The honest comparison
vLLM wins at high steady volume with a tuned rig.
Hosted wins for spiky, varied, or multi-provider workloads.
Model the TCO for your shape — don't trust averages.
Why agents still need a budget
Self-served inference doesn't govern the tools and APIs beyond it.
The firewall applies to any stack.
Where the money goes
| Cost bucket | How to control it |
|---|---|
| Hardware | Model size vs rig |
| Power | Utilization |
| Ops | Tuning time |
| Tools/APIs | Firewall rules |
FAQ
Is vLLM cheaper than hosted APIs?
At high steady utilization, often — model your TCO.
Does sipi.bot work with a vLLM stack?
Yes — the firewall governs tool and API spend regardless of where inference runs.
Related
Stop the next $12,400 night.
One API call (or MCP tool) in front of every agent transaction — APPROVED, BLOCKED, or FLAGGED, deterministic, ~5 ms, fully logged.
See plans — from $99/mo Try a live check