“We need model routing” is the request. What the person means is almost never clear, because four genuinely different capabilities share the word, and only two of them are configuration. The other two are engineering projects with prerequisites nobody mentions in the pitch.
The tell is what they say next. If it is “so we do not go down when the provider does”, they want failover. If it is “so we stop paying frontier prices for classification”, they want cost-based routing, and they are about to discover they need a way to tell whether the cheap model got it right. If it is “so each prompt goes to the best model for it”, they want quality-based routing, which cannot be adopted before you can score answers — the router is a prediction, and a prediction you cannot evaluate is a coin flip with a dashboard.
The second thing this category gets wrong is failover. Everyone writes the config; almost nobody works out what happens to a stream that dies at token 200, which is the failure that will actually reach users. That case has three possible behaviours, all visible to the user, and choosing between them is a product decision to make before the outage rather than during it.
Key takeaways
- Static failover and key-level load balancing are configuration. Cost-based and quality-based routing are engineering projects that require an evaluation loop first.
- A provider outage is rarely a clean 503. It is rising 429s, then climbing p99, then partial capacity — so health checks that look only for 5xx will keep routing into a brownout.
- A stream that fails at token 200 cannot be silently retried. You buffer and lose streaming, restart visibly, or continue with another model and accept a tonal seam. Pick one deliberately.
- Spreading traffic across providers lowers your cache hit rate. A cache keyed on the model means every reroute is a miss, and lost prefix-cache discounts can exceed the per-token saving.
- Together, Groq and Fireworks are places models run, not routers, even though they speak the same API shape.
Four things called routing
Static failover. An ordered list: try A, on failure try B. Pure configuration, and the only one of the four you can adopt in an afternoon. Its cost is hidden in behaviour drift, because B is a different model that will answer differently from the one your prompt was tuned against.
Load balancing across keys or regions. Spread requests across multiple API keys, accounts or regional endpoints to raise your effective rate limit and reduce the blast radius of one key being throttled. Also configuration, with one operational catch: your per-key quotas and your billing now live in several places, so cost attribution needs the router to record which key served each request.
Cost-based routing. Send the request to the cheapest model that will do the job. The hard word is “will”. Something has to decide whether a small model’s answer is acceptable, and there are only three mechanisms: a static rule keyed on request type, a cascade that tries the cheap model and escalates when a check fails, or a classifier over the prompt. The cascade is the honest one, and it needs a check — a schema validation, a confidence signal, a judge model — which is an evaluation problem wearing a routing hat.
Quality-based routing. Predict, per prompt, which model will answer best, and send it there. This is the only one that promises to be better rather than cheaper or more available, and it is the one you cannot adopt cold. Training or even validating a router means having labelled preferences over your own traffic, because a router tuned on public benchmarks is optimising for a task distribution that is not yours. If you cannot score answers today, build that first — the evaluation tools guide is the prerequisite, not the follow-up.
The practical consequence: adopt in that order. Failover, then key balancing, then a cascade with a real check, then prediction. Teams that skip to the end end up unable to explain why a given request went where it went.
What failover actually looks like during an outage
Start with the detection problem, because most failover configs assume a failure mode that does not happen.
A provider outage almost never arrives as a clean 503 on every request. It arrives as 429s multiplying while capacity tightens, then p99 latency climbing while p50 stays fine, then partial capacity where a third of requests succeed and the rest time out. A health check that looks for 5xx sees a healthy provider throughout. What you need instead is per-model, per-region tracking of error rate and a latency percentile, with a circuit breaker that opens on either — so you fail fast instead of paying the full timeout on every request before falling back.
Then the retry policy, where the standard mistake is to make things worse. A retry under load is extra load. Without a retry budget — a token bucket that permits retries only up to a small fraction of total requests — a provider brownout becomes you amplifying traffic into a struggling service and turning a degradation into an outage. Pair it with deadline propagation: only retry if the caller’s remaining time budget can absorb another attempt, otherwise fail now and say so.
Then idempotency. A completion is usually safe to retry twice, since the cost is tokens. A call that triggers a tool with side effects is not, which is why the agent side of this problem — idempotency keys derived from run and step identity — belongs in the agent frameworks guide. If your router sits in front of an agent loop, its retries must not re-execute tools.
The stream that dies at token 200
This is the case that decides how good your failover really is, and there is no clean answer.
You have already emitted 200 tokens to the browser. The connection to the provider drops, or the provider returns an error frame mid-stream. You cannot retry silently, because “silently” is exactly what is impossible — the user has read those tokens. Three options, and every product has to pick one:
Buffer server-side, emit on completion. Hold the whole response, so a mid-generation failure is invisible and retryable. You have also converted time-to-first-token into time-to-last-token, which for a long answer is the difference between a product that feels fast and one that feels broken. Defensible for short structured responses, wrong for chat.
Restart the response visibly. Clear what has been rendered, show a brief “regenerating” state, stream again from the new provider. Honest and simple. It is jarring mid-sentence, it doubles the wait, and if the user was already reading, they lose their place.
Continue with a different model. Send the 200 tokens back as an assistant prefix and ask the fallback to continue from there. Mechanically this works where the provider supports prefix continuation, and not every one does, so this option may simply be unavailable on your fallback. When it does work you get a tonal seam: a change of register mid-paragraph, a repeated transition, sometimes a contradiction of something in the first 200 tokens because the second model never reasoned its way there. Users notice. Whether they mind depends on whether your product is a chat assistant, where it reads as a glitch, or a long-form drafting tool, where it reads as an editing pass.
Two details that make all three worse if you skip them. Usage accounting: the failed stream consumed tokens you were billed for and your accounting probably never saw, because usage arrives in the final chunk and there was no final chunk. Attribute partial streams explicitly or your cost per request will quietly understate reality. And structured output: if the request demanded a JSON schema, a mid-stream failover to a model that does not enforce schemas produces plausible text that fails your parser. That is not graceful degradation, it is a different error later, and it is the argument for treating schema support as a hard filter on which models can serve as fallbacks at all.
Needs first-hand data: Put a proxy in front of your provider that kills the connection at a fixed token offset, then run each shortlisted router through it. Record, for each: whether the client saw an error, whether the response restarted or continued, whether token usage for the partial stream was reported, and whether a JSON-schema request survived the failover intact.
Sticky routing, cheapest routing, and what happens to your cache
Cheapest-per-token routing looks obviously correct on a spreadsheet and is often wrong once caching is in the picture.
Provider-side prompt caching discounts repeated prefixes — a long system prompt, a document, a few-shot block — and the discount is keyed to an exact prefix on a specific model at a specific provider. A router that spreads requests across three providers to chase per-token price splits that prefix three ways, so each provider sees a cold prefix more often and you pay full price on tokens that would otherwise have been discounted. On workloads with a large fixed prefix and short user turns, which describes most RAG and most agents, the lost cache discount can exceed the per-token saving that motivated the routing.
Your own semantic cache has the same problem in a different shape. Key the cache on model plus prompt and every reroute is a miss, so your hit rate falls exactly when a provider is degraded and you need the cache most. Key it on the prompt alone and you are serving one model’s answer to a request routed to another — which may be fine, or may quietly undo the reason you routed. That is a product decision, not a cache configuration, and it is worth reading alongside the semantic caching guide.
Sticky routing is the other side. Pin a conversation, a tenant or a session to one model, and you keep persona consistency, keep prefix caches warm, and get reproducible behaviour when someone reports a bug. You give up the ability to arbitrage price per request, and you need an explicit rule for what happens when the pinned model is unavailable — the answer is usually “fail over for this request, then return to the pin”, not “re-pin”, because re-pinning turns one bad minute into a permanently changed conversation.
The rule I would apply: pin by default, route on failure, and only arbitrage price on stateless, prefix-light request types like classification and extraction where no cache and no persona is at stake.
Inference providers are not routers
Together, Groq and Fireworks come up in every routing conversation because they expose the same request shape as everyone else, and swapping a base URL makes them feel interchangeable with a router. They are not the same category. They are places models run — infrastructure that hosts open-weight models and sells inference, each with its own hardware bet, and Groq in particular is a hardware story rather than a routing story.
A router’s job is to decide where a request goes and to hold the policy, the keys, the retry budget and the accounting. An inference provider’s job is to serve the request once it arrives. You put a router in front of inference providers, including these; you do not choose between them. Which providers you route to is a separate evaluation about model availability, latency profile and cost per token, and the API-compatibility questions that decide whether the swap is really a one-line change are covered in the OpenAI-compatible proxy guide.
OpenRouter

OpenRouter is a hosted aggregator: one API key and one endpoint reaching a large catalogue of models across many providers, with per-model fallbacks and provider preferences expressed in the request. It is the fastest way to make model choice a runtime parameter instead of a deploy, and its most underrated property is that for open-weight models served by several providers it can route between those providers, which is real availability insurance rather than just breadth.
Pros
- One key and one billing relationship instead of an account, a contract and a quota per provider
- Model and provider preference expressed per request, so trying a new model is a string change
- Multiple upstream providers for the same open-weight model gives genuine failover, not just a second model
Cons
- It is another hop in the request path, and a hosted one, so its availability becomes your availability
- Every prompt transits a third party, which is usually the first thing an enterprise review stops
- Feature support varies by upstream, so structured outputs, tool calling and prompt caching are not uniformly available across the catalogue
Best for: Teams that want breadth of model access and runtime model switching without negotiating with each provider, and whose data-handling review permits a hosted intermediary.
Pricing: Usage-based per token with a margin over the underlying provider rate, paid from a prepaid balance rather than per-provider invoices.
LiteLLM

LiteLLM is the default open-source answer, and it is two things: a Python SDK that normalises provider APIs, and a proxy server that exposes an OpenAI-compatible endpoint with routing, retries, fallbacks, virtual keys and budgets. Because you run it, the routing policy and every prompt stay inside your own network — which is why it appears in nearly every self-hosted design, and in the three-way gateway comparison.
Pros
- Self-hosted, so no prompt leaves infrastructure you control and no third party sits in the request path
- Routing strategies, retry and fallback policy, per-key budgets and rate limits are all configuration rather than code
- Very wide provider coverage, including self-hosted engines, so one config spans hosted APIs and your own GPUs
Cons
- You operate it: another service on the hot path for every model call, with its own scaling, upgrades and on-call
- Broad provider coverage means uneven depth; newer provider features arrive behind the provider’s own API
- The proxy becomes a single point of failure unless you run several instances behind a load balancer with shared state
Best for: Teams that need routing and spend controls without prompts leaving their network, and have a platform owner for one more service.
Pricing: Open source with no licence cost, plus a commercial enterprise tier for management and support features; the real cost is the infrastructure and the operator.
Portkey / Prisma AIRS AI Gateway

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks. The product is a gateway with routing at its centre — conditional routes, weighted load balancing, ordered fallbacks, retries with backoff, caching and per-key budgets, expressed as a config object rather than code. The corporate change matters for how you buy it rather than what it does: this is now an enterprise security purchase with a security vendor’s roadmap and procurement path, which helps if your blocker was a security review and adds friction if you wanted a self-serve tool.
Pros
- Routing as declarative config, including weighted splits and conditional rules on request metadata
- Guardrails, caching, budgets and observability in the same control plane as routing, so policy lives in one place
- Available as a hosted gateway or deployed into your own environment, so the data path is a choice
Cons
- Consolidation under a security vendor changes pricing power and roadmap priorities; assume enterprise packaging over self-serve
- The declarative config is expressive enough to become its own thing to learn, review and version
- Hosted mode puts a third party in the request path, which reopens the data-flow question you may have adopted a gateway to close
Best for: Enterprises that need routing, caching and guardrails under one governed control plane and prefer a security vendor as the counterparty.
Pricing: Tiered subscription with usage-based metering on requests, plus enterprise agreements for self-managed deployment; not itemised publicly at the enterprise tier.
Cloudflare AI Gateway

Cloudflare AI Gateway is a thin proxy at the edge: change your base URL, and you get caching, rate limiting, retries, fallbacks, logging and analytics across providers with no service to run. Its distinguishing property is where it sits — the request is already passing through Cloudflare’s network, so the added latency is small and the operational cost is close to zero. It is deliberately less configurable than a full gateway, and for a team whose problem is “we have no visibility and no fallback” that is the right amount.
Pros
- No infrastructure to operate; adoption is a base-URL change, which makes it the lowest-effort option here
- Caching, rate limiting and analytics arrive together, so one change closes several gaps at once
- Sits on a network your traffic may already traverse, keeping the added hop cheap
Cons
- Routing logic is simpler than a dedicated gateway: fine for fallbacks and retries, thin for conditional or weighted policy
- Hosted only, so prompts transit a third party and residency is governed by their controls rather than yours
- Ties an operational dependency to one vendor’s platform, which is a consideration if you are deliberately multi-cloud
Best for: Teams already on Cloudflare who want caching, fallbacks and per-request visibility this week without adding a service to run.
Pricing: Usage-based within the platform’s pricing model, metered on requests and logging volume rather than per seat, with an included allowance on lower tiers.
Vercel AI Gateway

Vercel AI Gateway targets the application developer rather than the platform team: one endpoint and one billing relationship across many models, with automatic failover, spend visibility and tight integration into the AI SDK that a lot of TypeScript applications already use. The pitch is that model access stops being a procurement task — no per-provider keys — and stays inside the framework you deploy with.
Pros
- Removes per-provider key and account management, which is the actual blocker for most small teams
- Works closely with a widely used TypeScript AI SDK, so routing is a parameter rather than an integration
- Automatic failover between providers and per-request spend visibility without configuration work
- Deploys with the application, so there is no separate gateway to run or scale
Cons
- Strong gravity toward one hosting platform; using it from elsewhere is possible but is not the design centre
- Less policy depth than a dedicated gateway — light on conditional routing, per-tenant budgets and guardrails
- Hosted intermediary, so prompt data flow and retention are governed by their terms, which you must verify
Best for: TypeScript product teams already deploying on Vercel who want multi-model access and failover without a platform project.
Pricing: Usage-based per token with a platform margin, billed alongside the rest of the platform rather than as separate provider invoices.
Not Diamond

Not Diamond is the clearest example of quality-based routing as a product: given a prompt, it predicts which model in your candidate set will answer it best and routes accordingly, with the option to train the router on your own preference data rather than generic benchmarks. That is a genuinely different value proposition from failover, and it comes with a genuinely different prerequisite — you need evaluation data, or you are trusting someone else’s notion of “best” for your task.
Pros
- Attacks quality per request rather than only cost or availability, which no configuration-based router can do
- Custom routers trained on your own preference data align the decision with your task instead of a public benchmark
- Often lands cost savings as a side effect, since many prompts genuinely do not need the largest model
Cons
- Requires labelled preference data to be worth anything; without an evaluation loop you cannot tell whether the router helps
- The routing decision is a model prediction, so it is probabilistic and harder to explain in a post-incident review than an ordered fallback list
- Adds a prediction step before the model call, which costs latency on every request
Best for: Teams with an existing evaluation pipeline and enough traffic diversity that per-prompt model choice is a real quality lever.
Pricing: Usage-based on routing decisions, separate from the model inference you still pay the underlying providers for.
RouteLLM

RouteLLM is a research-derived open-source framework rather than a managed service, and reading it is the fastest way to understand what quality routing actually does. The core idea is a learned binary decision — can a weaker, cheaper model handle this prompt, or does it need the strong one — trained on preference data, with a threshold you tune to trade cost against quality. Because it is a framework, it is a starting point you extend and operate rather than something you switch on.
Pros
- Open and inspectable, so the routing decision is a model you can retrain, threshold and audit yourself
- The strong-versus-weak framing maps directly onto the largest real saving available in most workloads
- An explicit cost-quality threshold makes the tradeoff a dial rather than a vendor’s opinion, and it runs in your own infrastructure
Cons
- Research-derived rather than a supported product: expect to do integration, retraining and operations yourself
- Router quality depends entirely on preference data representative of your traffic, which most teams do not have
- No gateway features around it — no budgets, no keys, no observability — so it is a component, not a platform
Best for: Teams with data science capacity who want to own a cost-quality router outright rather than pay for a decision they cannot inspect.
Pricing: Open source with no licence cost; you pay for the compute to train and serve the router plus the underlying model calls.
How to choose
Do this in order and stop when your actual problem is solved.
One: make failover work before anything else. Ordered fallbacks, a circuit breaker driven by error rate and latency percentile, a retry budget, and an explicit decision about the mid-stream case. Most teams asking for routing needed only this.
Two: decide where the router runs. If prompts may not transit a third party, that eliminates every hosted option immediately and you are choosing between self-hosted gateways — the constraint, not the feature set, is doing the deciding. Data-flow questions in depth are in the compliance gateway guide.
Three: measure your cache exposure before optimising per-token price. Work out what fraction of your input tokens are a repeated prefix. If it is high, pin by default and only arbitrage price on stateless request types.
Four: only then consider quality routing, and only if you can already score answers. If you cannot, the honest version of cost routing is a cascade with a hard check — schema validation or a judge — which gives you most of the saving and a signal you can act on.
| Tool | What it really does | Where it runs | Picks itself when |
|---|---|---|---|
| OpenRouter | Aggregation and provider-level failover | Hosted | Breadth and one billing relationship matter more than data path |
| LiteLLM | Self-hosted routing, keys and budgets | Yours | Prompts must not leave your network |
| Portkey / Prisma AIRS | Declarative routing plus guardrails and caching | Hosted or yours | You need one governed control plane and an enterprise counterparty |
| Cloudflare AI Gateway | Edge proxy with cache, retries, analytics | Hosted edge | You want visibility and fallbacks with nothing to operate |
| Vercel AI Gateway | Model access inside the app framework | Hosted | A TypeScript team wants multi-model without procurement |
| Not Diamond | Predictive per-prompt model choice | Hosted | You have evaluation data and quality is the lever |
| RouteLLM | Open cost-quality router framework | Yours | You want to own and retrain the routing decision |
Needs first-hand data: For one week of production traffic, log the repeated-prefix fraction of input tokens, the provider prompt-cache hit rate, and your semantic cache hit rate. Then replay the same week through a cheapest-per-token policy and compare total spend against sticky routing. The cache term is what decides it, and it is the number nobody has.
Frequently asked questions
Do I need a router if I only use one provider?
Yes, for two reasons that have nothing to do with model choice. Multiple keys or regions behind one endpoint raise your effective rate limit and contain the damage when one key is throttled. And a router is where per-tenant budgets, retry policy and per-request cost accounting live, which is the argument in the AI gateway hub and the cost tracking guide.
Will failover to another model break my prompts?
Assume yes until tested. Fallback models differ in instruction-following, tool-calling shape and structured-output enforcement, so your prompt is now running somewhere it was never tuned. Keep a small evaluation set that runs against every model in your fallback chain, and treat schema enforcement as a hard requirement for any model allowed to serve a structured request.
Should the router or the application own retries?
The router, so the policy is one place and the retry budget is global. An application that also retries turns one configured retry into several during an incident, which is how a degraded provider becomes an outage. Applications should own the deadline; the router should own the attempts within it.
Related reading
- Best AI gateways — the broader control plane that routing is one feature of.
- Best OpenAI-compatible proxies — where “compatible” breaks, which decides whether a failover is really a one-line change.
- Best semantic caching tools — the cache your routing policy is quietly undermining.
- Best LLM cost tracking tools — attributing spend per key, per tenant and per reroute.
- LiteLLM vs Portkey vs Kong AI Gateway — the three gateways compared on deployment model rather than feature grid.
- Best LLM evaluation tools — the prerequisite for any routing that claims to improve quality.