Guide to the Eval Gym
The eval gym is how sipi.bot proves the firewall works: 53 scenarios, 53/53 passing, machine-readable results. Here's what it actually tests.
What the scenarios cover
The six rule types against the documented failure shapes: retry loops, unknown vendors, category drift, overnight runs, key compromise, multi-tenant breaches.
Each scenario defines an agent state, a proposed transaction, and the expected decision.
Why a harness matters
A spend firewall is only as good as its edge cases. The gym encodes the incident database's patterns as tests.
Deterministic scenarios mean the results are reproducible — no model variance.
How to read the results
53/53 passing, published at /eval-report/ and machine-readable JSON at /eval.
None
| Scenario group | What it tests |
|---|---|
| Retry loops | Velocity limits |
| Unknown vendors | Merchant allowlists |
| Category drift | Category rules |
| Overnight runs | Time-of-day rules |
| Key compromise | Auth + caps |
FAQ
Are the scenarios real?
They encode documented incident patterns — sourced from the incident database.
Can I run them myself?
The eval results are public; the core is MIT-licensed, so you can run the engine and the scenarios.
Related
Stop the next $12,400 night.
One API call (or MCP tool) in front of every agent transaction — APPROVED, BLOCKED, or FLAGGED, deterministic, ~5 ms, fully logged.
See plans — from $99/mo Try a live check