One endpoint in front of every model — routed to the cheapest capable provider, failed over when one blinks, billed per call in USDC. No subscription, no markup — just cheaper, more reliable inference.
Auto-tiering per request, circuit breakers per provider, a failover receipt on every response. Typically 70%+ under baseline — and your own model stays yours; routing is opt-in.
Circuit breakers per provider, health tracking, and automatic failover to the next rung — with the full receipt on every response. Your agent never sees a provider outage; it just sees an answer.
Request model: auto and the gateway tiers,
costs, and fails over for you — pinned models stay pinned, and
local Ollama traffic always runs free.
Drop-in for anything that speaks OpenAI — Claude Code, your own agents, plain curl. Anthropic, OpenAI, OpenRouter, UsePod, and Nous/Hermes behind one interface, plus your local models for free. Keep your own model and API keys; routing is strictly opt-in.
No subscription, no seat fees, no markup. Every call settles in USDC over x402 at the provider's own rate, with spend ceilings per key so a runaway agent can't drain you. What you see in the ledger is what you paid.