No seats, no tiers, no markup. Every request settles in USDC over x402 at the provider's own rate — and routing picks the cheapest capable model, so the price you see is the price after the savings.
| Model | Tier | Input /1M | Output /1M | Context |
|---|---|---|---|---|
| gpt-4.1-nano | economy | $0.10 | $0.40 | 1M |
| claude-haiku-4-5 | economy | $0.80 | $4.00 | 200k |
| gpt-4.1-mini | standard | $0.40 | $1.60 | 1M |
| gpt-4.1 | frontier | $2.00 | $8.00 | 1M |
| claude-sonnet-4-5 | standard | $3.00 | $15.00 | 200k |
| local / ollama | free | $0.00 | $0.00 | your hardware |
The baseline for savings is a direct call to claude-sonnet-4-5 — what the same request would cost at provider list price, called direct. Request model: auto and the gateway tiers down when a cheaper model can handle the work; pin a model and it stays pinned.
Every API key carries its own spend breaker. When a key's recorded spend crosses its ceiling, it returns 402 with a top-up URL instead of quietly burning budget — and zero-cost local traffic never touches the breaker at all. Give each agent its own key so one drained session can't starve the others.
Each request settles independently — no float, no invoice, no "Stripe, coming soon." Watch every cent land in real time on the operator command center, or track your own key on the usage & ledger page.
Coming later: earn on the wait — sponsored placements on agent idle time. Advertiser? Get in early →