Gateway.
One endpoint for every model. Budgets, caching, fallbacks built in.
Gateway is the universal model proxy. Route every LLM call through one endpoint, attach budgets and rate limits to virtual keys, cache responses by semantic similarity, and fall back to a cheaper or alternate provider when the primary is down, over budget, or refuses. Every request becomes a perfect cost record that feeds Treasury, Router, and Reports without a per-provider SDK.
Route, cache, account.
One endpoint fronts every provider. Issue virtual keys with per-team budgets and rate limits.
Semantic caching deduplicates near-identical prompts. Automatic fallback when the primary refuses.
Every call captured with tokens and dollar cost. Feeds Treasury, Router, and Reports directly.
What it controls.
One URL, every model: Anthropic, OpenAI, Groq, Mistral, OpenRouter, local. Same SDK across all of them.
Per-team or per-project keys, each with its own monthly cap, rate limit, and access policy.
An identical request returns the stored answer, with the saving priced from the original call. Refuses to cache anything a stored answer could get wrong: temperature above zero, top_p below one, n above one, or tool calls.
Primary refuses or rate-limits, a configured fallback model continues the request seamlessly.
Every call captured with tokens and dollar cost. Feeds Treasury, Router, and Reports directly.
Who reaches for Gateway.
- ·Platform teams running shared AI infra
- ·Multi-provider AI orgs
- ·Cost-sensitive engineering teams
Pairs with the rest of Treasury.
Want Gateway in your stack?
We're onboarding design partners now. Join the waitlist to be in the Gateway cohort.
Just email is required. One email when Gateway goes live. Nothing else.