Engineering
What 4.1 billion routed requests taught us about model choice
Across a quarter of production traffic, 71% of prompts cleared the quality bar on a model costing a twentieth of the one they were sent to.
Read article
Engineering
Across a quarter of production traffic, 71% of prompts cleared the quality bar on a model costing a twentieth of the one they were sent to.
Read article
Release
Failover now completes in under 40 ms, the cache understands near-duplicate prompts, and two new regional endpoints are live.
Read article
Cost
Autocomplete fires on every keystroke. Sent to a reasoning model, it will quietly outspend everything else you have built.
Read article
Reliability
Retrying is easy. Retrying without changing the response shape, doubling the bill or breaking streaming is the hard part.
Read article
Evals
How we sample, score and discard real requests inside a retention window measured in hours.
Read article
Security
What actually crosses a border when a prompt is routed, and what the audit log records about it.
Read articleFree up to 50,000 requests a month. No card, no sales call, no migration.