Routing LLM traffic across inference providers by cost, speed and reliability
Providers now sell the same open-weight models at very different prices and speeds. We describe an adaptive router that keeps us on the cheapest provider that is fast and healthy, and moves our traffic when that changes, without an engineer involved.

The market for open-weight inference is heating up, and providers now sell the same models at very different prices and speeds. In this post we describe an adaptive router that keeps us on the cheapest provider that is fast and healthy, and moves our traffic when that changes, without an engineer involved.
Earlier this summer we wrote about moving most of our agent loops from Anthropic models to GLM, an open-weight model that several inference providers sell. At the end of that post we were sending all GLM traffic to Baseten and using Fireworks as a failover.
We have since replaced that setup with an adaptive router. It picks a provider for each request based on the cost and speed we measure in production, and it moves traffic away from a provider that starts returning errors. This post covers why we built it, how it works, the bugs we found after we turned it on, and what it has done for us. The numbers come from our production token ledger and logs.
The router does not handle every customer. Customers who require their data to be processed in a specific region still go to Anthropic or OpenAI models.
Fixed routers need babysitting#
Baseten, Fireworks and CoreWeave all serve GLM 5.2. They charge different prices, run at different speeds, and have outages at different times.
We started with round-robin. In our last full week on it, Fireworks served 51% of GLM 5.2 tasks and Baseten served 49%. Fireworks' prices were 25% higher than Baseten's and Baseten was faster, so half of our traffic cost more and took longer for no benefit.
We then switched to a fixed order, with Baseten first and Fireworks used only when Baseten failed. In the first full week on the fixed order, Baseten served 98.5% of tasks and Fireworks served 1.5%. That was cheaper and faster, but we still had three problems.
- Changing the order needed a code change and a deploy, and someone had to decide which provider was better that week.
- Our circuit breaker counted a 429 as a failure. Five 429s in a minute removed Baseten from the pool for the whole fleet, even when Baseten only wanted us to slow down a little.
- Outages required manually removing a provider from the pool to reduce the cost of failovers.
- A third provider, CoreWeave, had list prices about 45% lower than Baseten's. We had no data on how it performed, so we did not know where to put it in the order. We could have sampled it on a low percentage of traffic, but that again would have required a deployment.
We wanted routing that used the cheapest provider by default, took speed into account, moved traffic gradually when a provider returned rate-limit errors, and did all of this without a human stepping in.
How the router works today#
We'll first explain how the router works today, then we'll discuss the iterations it took to get here.
Everything is measured per task#
We measure cost and speed per task, not per model call. A task is one question or one code review, and it usually makes four to six model calls. We record every call in a ledger with the provider, token counts, duration, time to first token and task ID, and then add the calls up by task.
We do this because the task is what the customer pays for and waits for, and the cost of a task depends on more than the price per token. A provider that fails a call makes the task retry it, so the same task costs more tokens there. A provider with a worse cache hit rate bills more input tokens at full price for the same prompt. A provider that returns weaker output can make the agent loop take more turns. None of that shows up in the price of a single call. Speed works the same way: a customer feels the whole task, not one call inside it.
Every five minutes a job summarizes the ledger for each provider and task type and publishes the result to Redis. Each service keeps a copy in memory, so picking a provider does not make a network call. If the summary is too old to trust, for example because the sampling job stops working, the router uses the fixed order. We exclude tasks that used more than one provider, because we cannot assign their cost to either one.
Cost counts for more than speed#
Every provider serves the same model, so they differ in three ways: cost, speed and reliability. We handle reliability separately, which the next section explains. The score covers cost and speed.
textscore = 0.7 × (lowest cost / this provider's cost)
+ 0.3 × (fastest time / this provider's time)We try the provider with the highest score first, and the others in score order if it fails.
Why 0.7 and 0.3. The model is the same whichever provider serves it, so cost is the main reason to prefer one provider. Speed still matters, because people wait for these tasks. We chose the weights by deciding how much extra we would pay for speed. With 0.7 and 0.3, a faster provider can only come first if it costs less than 1.75 times the cheapest one, so we pay at most 75% more for speed.
In this example, B costs 10% more and is a third faster, and it comes first.
| Provider | Cost per task | Task time | Cost score | Speed score | Total |
|---|---|---|---|---|---|
| A | $0.10 | 30 s | 0.70 | 0.20 | 0.90 |
| B | $0.11 | 20 s | 0.64 | 0.30 | 0.94 |
Comparing apples to apples. A task takes longer when the answer is longer. We estimate how long each provider would take to write an answer of average length, using its time to first token and its tokens per second. Otherwise a provider that happened to serve longer answers would look slower than it is.
We also apply three rules after scoring.
- A provider that is more than three times slower than the fastest ranks last, whatever it costs. Cost is worth more than twice as much as speed, so without this rule a provider with a low enough cost than the alternative would come first however slow it was. A lower bound on speed, proportional to the fastest provider, is a simple rule that provides a minimum quality of service.
- To account for jitter, a provider must overtake the current first choice provider by 5% higher to replace it. The measurements change a little every five minutes. Without a margin, two providers with nearly equal scores would swap places on every update, and every swap sends traffic to a provider that has none of our prompts cached.
- A conversation stays pinned to the provider it started with for five minutes. Each turn of a conversation sends the whole history again. The provider that served the last turn has that history cached, so the next turn is cheaper and faster there. We use five minutes because that is roughly how long a prompt stays cached.
Errors are a gate, not a score#
Errors are not part of the score. If they were, a provider that was cheap enough would still rank first while it was failing.
Each provider has a permitted request rate, stored in Redis and shared by the fleet. The rate changes with each result.
| Result | Change to the permitted rate |
|---|---|
| 429 or 529 | multiplied by 0.5 |
| Other failure | multiplied by 0.8 |
| Success | plus 0.25 requests per second |
A 429 means the provider is asking for less traffic, so we halve the rate, as TCP does when it detects congestion. Other failures get a smaller cut, because the circuit breaker already handles a provider that is down. Recovery is slow on purpose, so a provider that has just failed takes minutes to get back to full traffic. This scheme is called additive increase and multiplicative decrease, and it converges to a stable rate whatever rate it starts from. When the first-choice provider is over its rate, the request goes to the next provider in score order. We never drop a request. The circuit breaker from the earlier post still handles a provider that is completely down.
Here is what that looks like in production. One evening Baseten returned two bursts of 429s about two hours apart. Within a few minutes of each one the router had cut Baseten's permitted rate and moved most requests to Fireworks. When the errors stopped, the rate climbed back and the traffic followed. No request failed for lack of a provider.
Every provider gets 5% of traffic, forever#
A provider that receives no traffic produces no measurements, so we would never find out that it had improved. We send at least 5% of requests to every provider in the pool. We chose 5% because the cost is small: at worst one request in twenty goes to a slower or more expensive provider. To mitigate customer impact, we recommend running parallel calls to the sampling pool from real work serving customers on the best provider. It's important to use real traffic so that sampling is reflective of the current state of customers.
We keep exploring a provider even after we have plenty of measurements for it. A provider's price and speed can change at any time, and the only way to notice is to keep sending it some traffic.
We built the off switch before we turned it on#
We first ran the router in production without acting on its output. For a week it logged what it would have chosen while the fixed order kept serving, and we compared the two.
In steady state it agreed with the fixed order, which was correct, because Baseten was better than Fireworks on both cost and speed. The useful signal was what it did when something went wrong. During that week Baseten returned bursts of rate-limit errors, and each time the router's backoff cut the traffic it would have sent there and then recovered as the errors cleared. Those incidents would have been resolved ahead of any intervention by the team, and after a week of sampling real traffic we were confident the router would work.
To mitigate the final risk of turning it on for customers we built an off switch first. A single Postgres row turns adaptive routing on or off. We change it from our admin console and every service applies the change within about five seconds. If the row is missing or cannot be read, routing uses the fixed order. We turned adaptive routing on in production a week after the shadow run, and it has been on since.
Initial rollout#
The first iteration wasn't completely smooth. Almost right away we encountered several issues that didn't present themselves while running in "shadow" mode.
One burst of errors could silence a provider for good#
Right away we discovered a bug that reduced the sampler to zero during a burst of errors without the ability to recover, which effectively took the provider out of the pool forever. To resolve this we still allowed the provider to be suppressed, but only to a minimum greater than zero, then switched to multiplicative increase to quickly bring it back to the minimum sampling rate.
The leader looked cheaper because it was the leader#
Providers charge less for prompt tokens they have cached, and they return the first token sooner. The provider serving most of the traffic has the warmest cache, so its measured cost and latency are lower than they would be otherwise. A provider that only receives exploration requests always has a cold cache. Comparing them on measured numbers favours whichever provider is already first.
To remove the bias, for ranking only we price every cached token at the uncached rate, and we measure time to first token only on calls that read nothing from cache. That puts every provider on the same footing, but it also throws away real information, because providers do not cache equally well.
So we brought caching back in without the bias. Each provider gets a cache hit rate taken from the last time it was in steady state, meaning a period when it carried most of the traffic and its cache was warm. That rate is weighted by how recently that period was, so a provider that led last week is scored on last week's hit rate, and one that has not led for a month is scored closer to the pool average. A provider that has never led is scored at the pool average until it does.
Slow providers looked healthy once traffic dropped#
Speed uses the 30-minute window when it has enough samples. When a provider slowed down, it lost traffic, its 30-minute sample became too small, and the router fell back to a 48-hour average. Most of that average came from before the slowdown, so the signal that should have moved traffic was buried under healthier numbers from the day before.
We now use a small recent sample when it shows the provider got slower, and ignore it when it shows the provider got faster. This means we are slow to send traffic back after a provider recovers. We accept that, because missing a slowdown is worse.
Provider history was wiped every hour#
When we add a provider, we do not want to send it all of our traffic at once. We have no record of how it behaves under load, and a sudden jump could overload it. So the router limits a provider that has no recent history to a small share of traffic, and raises the limit as successful requests come in.
A bug in how we stored that history cleared it at the start of every clock hour. Once an hour, Baseten looked like a provider we had never used, even though it was serving almost all of our traffic. The router limited it, and the requests it could no longer take went to Fireworks. Within a minute Baseten had enough new history and the limit came off.
We could see this in the ledger. In the week before the fix, Fireworks served 22.9% of GLM 5.2 calls in the first minute of each hour, compared with 4.6% for the rest of the hour.
We fixed it in two ways. History now expires gradually, one minute at a time, so it is never empty for a provider that is in use. The slow start also now applies only to a provider we have never used. In the week after the fix, the first minute was 11.8% compared with 7.2% for the rest of the hour. That is better, and we have not yet found the cause of the rest.
Not every error is the provider's fault#
Code review sometimes uses its whole output token budget. We counted that as a provider failure, which lowered the provider's rate and retried the review on another provider. The retry ran out of budget too, and we paid for both. In one three-day audit this was nearly every failure recorded for Baseten and Fireworks. We no longer count it against the provider.
Some requests can only be tried once, because they include reasoning state that only the original provider can continue. Exploration was using those requests, so when an exploration request went to a provider that was down, the user got an error. In three days, 71% of the requests that ran out of retries were requests that could not fail over. Now we always send those requests to the first-choice provider and never use them for exploration.
We also reduced the rate once for every rejected request. Providers reject in bursts, so a busy provider with a 5% error rate dropped to its minimum rate in about 30 seconds. We now reduce at most once every five seconds, so that one burst counts as one event. We also only reduce when the recent error rate is above 10%, so that occasional errors do not move traffic. TCP NewReno does the same: it treats several packet losses in one window as a single event and reduces its window once.
What the router has done for us#
In its first four weeks the router saved our necks seven times, and no engineer had to do anything for any of them.
- It found a provider that was more than twice as fast as the one we were using.
- When that provider buckled under the full load a day later, the router pulled traffic back to Baseten and returned it once the provider recovered.
- Our account with that provider ran out of credit, and the router backed traffic off it in minutes.
- Four times in one week, Baseten slowed down, and the router moved traffic to Fireworks and back again.
With the fixed order, each of these would have needed someone to notice the problem and intervene. During the slowdowns Baseten was still returning successful responses, so the circuit breaker would not have reacted either, and customers would have waited three to five times longer for a response to start.
The rest of the time the router makes the same choice the fixed order did, which is Baseten first with 5% of requests going to Fireworks.
It found a provider we had never used#
CoreWeave had never served production traffic. It received its first exploration requests within an hour of the router going live. The router promoted it because its cost estimate said CoreWeave was about half the price of Baseten, and it was returning the first token in about 3 seconds against Baseten's 7 to 9. Five hours after its first request it was serving 92% of GLM 5.2 tasks.
The next day it served 64%, and the gap is the router doing its job. Late that morning CoreWeave started to struggle under the full load. Its time to first token went from 3 seconds to between 50 and 170 seconds, and its success rate fell from 98% to as low as 50%. The router demoted it on both counts: the latency rule ranked it last for being far slower than Baseten, and the rate limit cut the traffic it was allowed to serve as the errors came in. For about eight hours most traffic went back to Baseten. When CoreWeave recovered that evening, its time to first token back at 3 seconds, the router promoted it again and it served 90% of tasks for the next three hours.
Here is what we actually paid and waited for, on tasks served entirely by one provider.
| Paid per task | Median time | p90 time | |
|---|---|---|---|
| Search, Baseten, fixed order | $0.0603 | 29.4 s | 88.2 s |
| Search, CoreWeave, day 2 as leader | $0.0611 | 13.1 s | 46.2 s |
| Code review, Baseten, fixed order | $0.1188 | 100.4 s | 271.3 s |
| Code review, CoreWeave, second stint as leader | $0.0909 | 44.8 s | 150.1 s |
The speed was real. Median task time fell by 55% for both task types. The cost was not what the router thought. We paid 23% less for code review and the same for search, against an estimate of about half. At the time, the router priced every provider as if nothing were cached, and CoreWeave caches far less of our prompts than Baseten does, so its discount only applied to the tokens it missed. This has since been resolved. See the discussion on caching above.
It handled an outage we caused ourselves#
Two days after CoreWeave took the lead, our account with it ran out of prepaid credit and every request returned a 402. In the ten minutes before that, CoreWeave served 95% of GLM 5.2 calls. In the minutes following, CoreWeave was removed from the pool and Baseten took over the traffic. No engineer touched the configuration.
As soon as we added credit, sampling resumed and CoreWeave was restored as the leader.
This is the promise of adaptive routing fully realized. Use the cheapest, fastest option until it is exhausted, then fail over to the next best provider without anyone noticing.
It rode out four slowdowns on Baseten#
We added GLM 5.3 to the router the day after going live, with a configuration change. Baseten and Fireworks serve it, and their costs are within 3% of each other, so the ranking depends on speed.
Baseten normally returns the first token for a GLM 5.3 search in about 5 seconds. Four times in one week that rose to between 14 and 26 seconds, and each time the router moved most GLM 5.3 traffic to Fireworks until Baseten recovered.
| Slowdown | Baseten time to first token | Lowest hourly share of calls on Baseten |
|---|---|---|
| First | 15 to 21 s | 32% |
| Second | 16 s | 14% |
| Third | 25 to 26 s | 25% |
| Fourth | 14 to 21 s | 5% |
Here is the fourth one in detail. Fireworks served 4.5% of GLM 5.3 calls in the hour before it started, 81.5% in the next hour and 95.5% in the hour after that. When Baseten recovered, Fireworks went back to 11.9% and then 4.1%. Baseten's success rate stayed above 95% during all four slowdowns, so the circuit breaker would not have reacted.
What it did to the bill#
We want to be straight about cost, because the summary of this post is about spend. Here is GLM 5.2 cost per task by week, priced from each provider's rate card. The table stops where most of our traffic moved to GLM 5.3, because the tasks left on 5.2 after that are a different mix and the weeks stop being comparable.
| Week | Routing | Search, cost per task | Code review, cost per task |
|---|---|---|---|
| 1 | round-robin | $0.070 | $0.128 |
| 2 | fixed order, Baseten first | $0.060 | $0.123 |
| 3 | adaptive from midweek | $0.065 | $0.117 |
| 4 | adaptive | $0.063 | $0.120 |
The drop came from leaving round-robin: search fell 14% and code review 4% in the first week on the fixed order. The router did not lower cost further, because Baseten was already the cheapest healthy provider most of the time. CoreWeave's price advantage never had the chance to show up in a weekly number: it hit capacity problems under our full load, and we had to take it out of the pool before its lower prices could work through the bill. Those capacity problems have since been resolved, and CoreWeave is back in the pool. What the router bought us was staying at that floor without a person watching, at a cost of about 1% for exploration, while median task time fell by half whenever a faster provider was available.
The table also leaves out the cost we did not pay. With a fixed order, every one of Baseten's slowdowns would have landed on customers. During the worst of them, Baseten took 25 seconds to return the first token instead of 5, and a fixed order would have kept sending it every request until an engineer noticed and shipped a change. The router sent that traffic to Fireworks, which costs about 3% more for GLM 5.3. Paying a few percent more for an hour so that customers do not wait five times longer is a trade we will take every time.
All this being said, we do expect costs to fall dramatically in the future. Most of our traffic is now on GLM 5.3, which CoreWeave does not serve yet. Once it does, a provider with prices 45% below Baseten's will be in the pool for our largest workload, and the router will move traffic to whichever provider is cheapest that week. Providers can now compete for our traffic on price, and the winner changes without a deploy.
What is still wrong with it#
All services switch provider at the same time. They read the same summary and pick the same top score. When load information is out of date, sending every task to the provider that looks least loaded can hurt performance. The rate limit reduces the effect after it starts.
We always pick the top score. Thompson sampling draws a random estimate for each provider from what has been measured so far, and picks the provider with the best draw. That would stop all services switching at once, and we would not need a fixed exploration rate. We kept the simpler rule because we can reproduce any routing decision from the saved statistics, which is how we found every problem in this post.
Advice for anyone building this#
- Keep errors out of the score and handle them with a separate rate limit. If errors are a term in the score, a low enough price cancels them out.
- The router's choices change the data it measures. Correct for that, and check the correction against what you were billed.
- Keep exploration on permanently, pick at random when providers are tied, and do not explore with requests that cannot fail over.
- Let a small recent sample lower a provider's ranking, and require a full sample before raising it.
- Build the off switch before you turn it on, and save the statistics used for every routing decision. We needed both in the first week.
Open-weight providers are turning inference into a commodity. Several of them sell the same model, and what separates them is price, speed and reliability, all of which change from week to week. A router that adapts in real time is what turns that into a benefit.
In its first four weeks ours moved traffic seven times without an engineer involved, including a provider outage where traffic was off the dead provider in minutes. Median task time fell by 55% whenever a faster provider was in the pool. Cost per task stayed 14% below what round-robin had been costing us, and it stayed there without a single deploy. We expect the bigger savings are still ahead. New providers keep entering the market and supply keeps growing, and an adaptive router is what lets us take the full benefit of every price cut the moment it happens.


