Home / Home / Guides / Guide to the Eval Gym

Guide to the Eval Gym

The eval gym is how sipi.bot proves the firewall works: 53 scenarios, 53/53 passing, machine-readable results. Here's what it actually tests.

What the scenarios cover

The six rule types against the documented failure shapes: retry loops, unknown vendors, category drift, overnight runs, key compromise, multi-tenant breaches.

Each scenario defines an agent state, a proposed transaction, and the expected decision.

Why a harness matters

A spend firewall is only as good as its edge cases. The gym encodes the incident database's patterns as tests.

Deterministic scenarios mean the results are reproducible — no model variance.

How to read the results

53/53 passing, published at /eval-report/ and machine-readable JSON at /eval.

None

Scenario groupWhat it tests
Retry loopsVelocity limits
Unknown vendorsMerchant allowlists
Category driftCategory rules
Overnight runsTime-of-day rules
Key compromiseAuth + caps

FAQ

Are the scenarios real?

They encode documented incident patterns — sourced from the incident database.

Can I run them myself?

The eval results are public; the core is MIT-licensed, so you can run the engine and the scenarios.

Related

Stop the next $12,400 night.

One API call (or MCP tool) in front of every agent transaction — APPROVED, BLOCKED, or FLAGGED, deterministic, ~5 ms, fully logged.

See plans — from $99/mo Try a live check