Our story
One hard-coded model name, and a bill nobody could explain.
In 2023 our founding team shipped a support assistant. It worked, users liked it, and three months later finance asked why a feature used by 4% of customers accounted for 60% of the cloud spend.
The answer was boring: every request — including "where is my order" — went to the largest model available, because choosing per request meant writing a router, and nobody had time to write a router.
So we wrote one. It cut the bill by half in a week. Then a provider had an outage and the router handled that too, and it became obvious this was not a feature of one product. It was infrastructure that every team building on models was going to need, and most of them were going to build it badly, twice.
- 2023
An internal script
Two hundred lines of Python deciding which of three models got the request. Deployed on a Thursday.
- 2024
Spun out
Six design partners, all of whom had written some version of the same script. Failover and spend ceilings landed that year.
- 2025
Evals and residency
Nightly grading on live traffic, regional pinning, and the audit log that let the first regulated customers sign.
- 2026
Twelve providers
4.1 billion requests routed in the last quarter across 2,400 workspaces, at 31 ms of added latency.
A small company holding a large amount of traffic.
Four rules we have not broken yet.
We are not precious about the roadmap, but these four hold — even when a customer asks us nicely to bend one.
0
Exceptions made, across three funding rounds and one very persuasive enterprise RFP.
Never mark up inference
You keep your own provider keys. If we profited from the models you use, you could not trust the routing decision.
Ship the boring thing first
Failover, logging and spend caps before anything with a demo video. Infrastructure earns trust by being dull.
Default to dropping data
Prompt bodies are discarded unless you turn retention on. The safest place for your customers' text is not our disk.
Support is engineers
The person answering in Slack can read the router source. There is no first line, and no ticket queue to escalate out of.
The people who answer when you write in.
Thirty-eight of us across nine countries. We are hiring in platform engineering and developer relations.
Route your first request in about four minutes.
Free up to 50,000 requests a month. No card, no sales call, no migration.