LiteLLM is usually adopted in an afternoon and questioned about nine months later. The afternoon goes well: one container, a YAML file, an OpenAI-compatible endpoint, every provider your team wanted behind one key. The nine-months-later conversation is different, and it is almost never about features.
It is about a Postgres nobody wants to own, a pinned version three releases behind because the last upgrade broke a config key, a pod that got OOM-killed during a demo, and the growing awareness that every AI feature in the product depends on a service whose only maintainer is now on another team.
That is not a criticism of LiteLLM. It is what happens when you self-host a critical-path service without deciding who owns it, and the same story exists for every open-source proxy in every category. The tool is fine. The operating model was never decided.
So the useful question is not “what is better than LiteLLM”. It is which of four specific things you are trying to get away from, because each points at a different shortlist and three of the four have nothing to do with routing quality.
Key takeaways
- Operational load, missing enterprise governance, control-plane duplication and edge placement are four different reasons with four different answers.
- What you inherit when you self-host is a proxy tier, a Redis, a Postgres full of prompt payloads, and the uptime of every AI feature you ship.
- Python in a hot path is a sizing and backpressure concern rather than a disqualifier — but it is one you have to actively manage, especially with many concurrent long-lived streams.
- LiteLLM still beats every alternative here on provider coverage, config-as-code, and having no vendor in the data path. If those are why you chose it, fix the operating model instead of migrating.
What you inherit when you self-host a proxy
A self-hosted LiteLLM deployment is three components, and only the first behaves.
The proxy tier is stateless, horizontally scalable and disposable. Redis holds shared rate-limit and budget counters plus the response cache — both correctness requirements once you have more than one replica, because a limit enforced in each pod’s local memory is your limit multiplied by your replica count, producing no error, just a bill. Postgres holds virtual keys, teams, budgets, the spend ledger, model configuration and, if enabled, the full text of every prompt and completion that passed through. That last part turns a storage decision into a data-retention decision: prompt logs need a defined lifetime, an enforcement mechanism, an access policy stricter than most production tables, and an indexed tenant identifier so you can answer an erasure request.
On top of that you inherit the operational surface: being in the request path, so two replicas minimum and readiness probes that test the database and cache rather than the listening socket; streaming restarts, where a pod that takes SIGTERM and exits promptly hands the client a truncated answer mid-sentence with no error frame; retry policy spread across four layers — application, SDK, gateway, fallback chain — which multiply a thirty-second provider degradation into several times normal upstream load; and upgrades on a fast-moving project with a growing config surface.
The self-hosted gateway operations piece goes through all of this in detail. If that list makes you think “we have not done four of those”, the honest first move is to fix the deployment rather than change the product, because every self-hosted alternative here has the same list.
Python in a hot path, described accurately
This concern is stated badly in both directions — either “Python is too slow for a proxy” or “it makes no difference, it is all I/O”. Both are wrong, and the accurate version is specific enough to act on.
A gateway is I/O-bound work: accept a request, hold a connection open, forward bytes, stream bytes back. Async Python handles I/O-bound concurrency well, and for a request that spends multiple seconds waiting on a model the interpreter is not the bottleneck. That much of the defence is correct.
What is genuinely different is that a single process runs one event loop, and anything that blocks it stalls every concurrent request in that process, not just its own. The failure modes that follow:
- CPU work on the event loop. Tokenisation for token counting, JSON serialisation of large payloads, cache-key hashing, cost arithmetic, inline guardrail checks. Individually small; on a large body, in front of hundreds of concurrent connections, they add latency to unrelated requests. This is the mechanism behind “our p99 got worse and the models were fine.”
- Worker sizing versus connection count. You scale by running multiple worker processes, each with its own event loop, memory footprint and upstream pools. Too few and CPU work in one loop hurts everything on it; too many and per-worker memory plus per-worker pools multiply into resource pressure and provider-side connection limits. There is no universal number — it is a function of payload sizes, streaming concurrency and instance shape.
- Long-lived streams occupy a slot for their whole duration. Concurrency here is not requests per second, it is concurrent open streams: a much smaller number that behaves very differently, and the one your capacity planning should use.
- Backpressure. With a slow consumer the proxy either buffers, growing memory per stream, or applies backpressure to the upstream read, which is what you want. Verify which happens, because unbounded per-stream buffering is how a proxy tier OOMs during a traffic spike rather than during a load test.
None of this makes Python the wrong choice. It makes the proxy tier something you load-test with your payload sizes and your streaming concurrency rather than with a request-per-second benchmark that tells you nothing. Alternatives written in C++, Go or on an edge runtime avoid the event-loop-blocking class of problem structurally, which is a real advantage and not the same as being faster.
Needs first-hand data: Load-test on concurrent open streams rather than requests per second. Hold N streaming generations open simultaneously, stepping N up until p99 time-to-first-token degrades, and record resident memory per worker at each step. Repeat with a deliberately slow consumer reading one byte at a time. Those two curves are your real capacity model and neither appears in any documentation.
Reason one: you do not want to operate it
The most common reason and the most honest one. Running the proxy, its Postgres and its Redis, and being in the request path for every AI feature, is a platform-team job. If you do not have a platform team, you have a volunteer, and volunteers change teams.
The tell is not an outage. It is the version you are pinned to, the alert nobody tuned, and the fact that the person who wrote the YAML now works on something else. A gateway you cannot confidently upgrade is a liability regardless of how good it is.
Moving to a managed gateway trades that for a vendor relationship, a data path that includes their infrastructure, and a bill. For a small team shipping product features that trade is usually correct — the incident where every AI feature goes down because of your own infrastructure is more expensive than the subscription.
Shortlist: Portkey / Prisma AIRS for the full managed enterprise gateway, Cloudflare AI Gateway or Vercel AI Gateway if you are already on those platforms and your needs are modest, OpenRouter if you want model access and someone else’s operational problem, and TrueFoundry if you need it managed and inside your own cloud account.
One caution: if your reason for self-hosting was data control, a managed gateway does not solve reason one without reopening that question. The requirement has not gone away because operating it got tiring — the Portkey alternatives piece is that same tension from the other direction.
Reason two: you need enterprise governance out of the box
This shows up as a security review you cannot pass or a procurement conversation you cannot finish.
The asks are always the same. SSO against your identity provider, so gateway access is not a shared key in a password manager. Role-based access control with real roles: who can mint a virtual key, who can raise a budget, who can read prompt contents in logs. Audit trails — not request logs, but a record of who changed which configuration when, exportable. Budget enforcement, meaning the gateway refuses the request that would breach the limit rather than showing overspend afterwards. And increasingly guardrail policy applied centrally rather than in every service.
You can assemble most of this around a self-hosted proxy: an identity-aware proxy in front for SSO, your own audit pipeline from the database’s change log, budgets enforced by the gateway’s limits, guardrails as a service you call. Teams do it and it works. What it does not do is compress into a security questionnaire answer, and the assembly is undocumented internal glue that one person understands.
Note the specific trap: some of LiteLLM’s deeper governance sits behind its paid enterprise tier. “We need governance, therefore we must leave the open-source proxy” is sometimes a false step — the comparison you want is LiteLLM’s enterprise tier against the managed platforms, not the open-source build against them.
Shortlist: Portkey / Prisma AIRS, TrueFoundry where the deployment must stay in your VPC, and Kong AI Gateway if your enterprise governance already exists in Kong and you would rather extend it than buy a second one. The enterprise gateway requirements piece itemises what auditors actually ask for.
Reason three: you want one gateway of record
If your organisation already runs an API gateway, a self-hosted LiteLLM is a second control plane, and that is a cost you pay continuously rather than once.
Your REST traffic is governed by the gateway: authentication, authorisation, rate limits, logging destinations, WAF rules, network policy, a service catalogue, and a team that owns all of it. Your LLM traffic is governed by LiteLLM, with its own key namespace, policy model, logs and notion of a team. Two consoles, two change processes, two answers to any audit question, and two identity models that disagree in some subtle way you find during an incident.
The case for consolidating strengthens as agent traffic grows. Once agents call MCP servers, MCP servers call internal APIs, and agents call other agents, the graph looks like a service-mesh problem, and it becomes difficult to argue the component governing service traffic should not govern this too.
Shortlist: Kong AI Gateway if Kong is your gateway, since it explicitly frames itself around governing LLM, MCP and agent-to-agent traffic through one control plane; Envoy AI Gateway if Envoy is your data plane and you want LLM routing expressed as Kubernetes resources alongside the rest of your ingress.
What you trade is LLM-specific depth. Provider coverage, prompt handling and per-request cost attribution are better in a product whose entire job is model traffic. If the audit question is real in your organisation, one policy layer is worth more than the depth.
Reason four: you want the gateway at the edge
Narrow, and where it applies, decisive.
A gateway in your cluster adds a hop from the caller to that cluster. If the caller is a browser or a mobile client and the cluster is in one region, that is a real round trip added to every model call before the model has done anything. Deploy the gateway in one cluster and serve applications in three regions and you have quietly added cross-region latency to a user-facing path — a common accident, because the gateway gets deployed wherever the platform team’s cluster happens to live.
An edge-deployed gateway inverts that: the proxy runs near the caller, and the long leg is the one you cannot avoid anyway. For time-to-first-token on an interactive streaming feature, that is the difference between a gateway being free and a gateway being noticeable.
Two checks before treating this as your reason. Whether your callers are actually geographically distributed — if all traffic originates from services in the gateway’s region, edge placement buys you nothing. And whether your caching survives at the edge, because a cache spread across many locations has a lower hit rate per location than one central cache with the same total traffic.
Shortlist: Cloudflare AI Gateway, Vercel AI Gateway, and Portkey / Prisma AIRS, whose gateway has a lightweight edge-deployable heritage.
Needs first-hand data: Measure time-to-first-token from each region your users are in, direct to provider versus through your current in-cluster gateway. If the delta is a meaningful fraction of TTFT for any region, edge placement is a real reason. If it is a rounding error everywhere, cross reason four off.
OpenAI-compatibility: the standard LiteLLM sells, not a product
Worth separating explicitly, because it explains both LiteLLM’s main asset and why leaving is tractable.
The OpenAI request and response shape — a messages array, tool definitions, an SSE stream of deltas at a /v1/chat/completions-style path — became the de facto wire format for talking to a language model. LiteLLM’s core value is a very large library of mappings from that shape onto every provider’s actual API, plus the reverse translation on the way back. The format itself is a specification with no vendor, no dashboard and no bill; what LiteLLM sells is coverage of it.
What it gives you
- One client library and one request shape across hosted providers and your own inference servers
- Gateway migration is a base-URL and key change, because every option here presents the same front door
- You can run two gateways side by side and shadow traffic to compare them honestly
- Application-level code survives the gateway decision entirely, which is why this migration is cheaper than an observability migration
What it does not do
- Compatibility is partial and uneven. Provider-specific parameters, reasoning controls, cache hints, safety settings and multimodal payloads are each handled differently, and that is where a migration breaks
- Error semantics are not standardised, so how a gateway normalises rate limits, retry-after headers and content filters determines whether your retry logic is correct
- It carries none of your configuration — routing rules, key hierarchies, budgets, guardrails and prompt templates all move by hand
- Token accounting differs by provider and tokeniser, so every gateway’s unified cost number is an approximation it assembles
Test the specific parameters your product depends on against each provider through each candidate. The OpenAI-compatible proxy landscape goes further into where the abstraction leaks.
What LiteLLM does better than everything below
Say this plainly, because a reader who leaves for the wrong reason comes back in six months having spent a quarter on it.
Provider and model coverage. Nothing else here is close. In a market where new models and endpoints appear constantly, the breadth of LiteLLM’s mapping library is its real moat, and a managed gateway supporting the top handful of providers will eventually block you on a model you want.
Configuration as code, in your repository. Model lists, fallback chains, load-balanced deployment groups and budgets are a YAML file under review, not console state someone changed on a Friday. Every managed alternative moves at least some of that into a UI, and UI configuration has no diff, no reviewer and no rollback.
No vendor in the data path, and no exit cost. Prompts go from your service to your proxy to the provider — a claim no managed gateway can make, and why LiteLLM remains the default for residency and air-gapped deployments. There is also no contract, no export request and no history stranded in someone else’s schema.
Self-hosted inference servers as first-class citizens. Pointing at your own vLLM deployment is another entry in the same config as a hosted provider. Managed gateways range from awkward to impossible here.
If two or more of those are why you adopted it, the correct move is almost certainly to fix the operating model — name an owner, get to two replicas, sort out retention, fix the retry layering — rather than to migrate. Reason one is a staffing problem wearing a technology costume, and buying a managed gateway to solve it works only if you were also willing to give up the four things above.
Portkey

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks, generally available for enterprises, and it is the most direct managed answer to reasons one and two. The product is a managed gateway — unified provider API, configurable routing and fallback, virtual keys, caching, guardrails, prompt management and request-level observability — whose centre of gravity has moved to security and governance under its new ownership. The gateway core has an open-source lineage and a lightweight, edge-deployable design, which also makes it relevant to reason four.
Pros
- Removes the three-tier stateful service from your team entirely, which is the whole point if operational load is your reason
- The governance surface enterprises get blocked on — SSO, RBAC, audit trails, budget enforcement, guardrail policy — is product rather than assembly
- Security and AI runtime protection are the organising idea, an advantage where a security team owns AI risk
- Lightweight, edge-deployable heritage, so the hop can sit near the caller rather than in one distant cluster
Cons
- You are buying into a security vendor’s platform, with the roadmap, packaging and procurement posture that implies
- Prompts traverse vendor infrastructure on the managed path, so it does not satisfy the data-control requirement that makes teams self-host
- Enterprise motion rather than self-serve for the full product, so evaluation is a conversation
- Provider coverage and config-as-code are both weaker than what you are leaving
Best for: Teams that need an enterprise governance surface and a gateway they do not operate, where a Palo Alto Networks relationship is an asset rather than a complication.
Pricing: Commercial, sold on enterprise agreements with usage-based metering on gateway traffic rather than a single public list price.
Kong AI Gateway

Kong AI Gateway is Kong Gateway plus a family of AI plugins, and it is the answer to reason three: one gateway of record instead of two control planes. Its product framing covers LLM, MCP and agent-to-agent traffic through the same gateway. Every plugin you already run for authentication, authorisation, logging and rate limiting applies to model routes unchanged, and it deploys self-hosted — so it addresses control-plane duplication without necessarily addressing operational load.
Pros
- One policy layer and one team for REST, LLM, MCP and agent traffic, which removes the duplicate-governance tax permanently
- Existing authentication, rate-limiting and logging plugins apply to model routes, so nothing is rebuilt in a second system
- Deployable self-hosted or with a managed control plane, so the operating model is a choice
- Positioned for agent and MCP traffic rather than only model calls, which is where the governance problem is heading
Cons
- Does not reduce operational load if you self-host it — you have swapped which service you operate, not whether
- The advantage is conditional on already running Kong; adopting it for LLM traffic alone is a lot of gateway for a narrow job
- The open-source and enterprise split across the AI plugin set must be verified early, because an open-source evaluation may not match what you ship
- LLM-specific depth — provider coverage, prompt management, cost attribution — trails the specialists
Best for: Platform teams already running Kong who want to collapse two control planes into one and are heading towards agent and MCP traffic they will have to govern.
Pricing: Open-source gateway with no licence cost plus an enterprise subscription for the advanced plugin set and managed control plane.
Cloudflare AI Gateway

Cloudflare AI Gateway addresses reasons one and four together, an unusually good fit for a team whose complaint is “we do not want to run this and our users are everywhere”. It sits on Cloudflare’s edge network, you adopt it with a base-URL change, and you get caching, rate limiting, retries and fallback, request logging, analytics and cost visibility with nothing to operate.
Pros
- Nothing to operate: no proxy tier, no Redis, no Postgres, no retention policy for prompt logs in your own database
- Edge placement means the hop does not add a cross-region round trip for distributed callers
- Adoption is a base-URL change, and removal is the same change in reverse, so it is a low-risk step
- Usually a configuration change inside a platform you already have rather than a new vendor and a new security review
Cons
- Managed only, with no self-hosted or air-gapped path, so it reverses the data-control decision entirely
- Governance is well short of an enterprise gateway: team hierarchies, granular RBAC, prompt versioning and audit trails are not the product
- Provider and model coverage is narrower than LiteLLM’s, and a self-hosted inference server is a less natural fit
- Deepens concentration on one platform vendor, a different version of the dependency you were managing
Best for: Teams already on Cloudflare whose requirements are one endpoint, fallback, caching and a cost number, with distributed callers and no residency constraint.
Pricing: Usage-based within the Cloudflare developer platform, metered on gateway traffic and logging rather than an enterprise gateway licence.
Vercel AI Gateway

Vercel AI Gateway is the reason-one-and-four option for teams whose application already lives on Vercel. One endpoint and one key across many models, spend limits, provider failover and usage observability, wired into the SDK the team already calls models through. Its appeal is the same as the Cloudflare option: a configuration change inside a platform you already pay for rather than a new operational commitment.
Pros
- Close to zero integration cost for a team already building on Vercel, since the gateway is part of the toolchain they use
- One key across many models removes provider account sprawl without a procurement project
- Spend limits and usage visibility cover the cost controls most teams actually exercise
- No infrastructure, no database, no retention policy, no on-call
Cons
- Managed only, so it fails any residency or air-gap requirement outright
- Governance is thin against both LiteLLM’s enterprise tier and a dedicated enterprise gateway
- Tightly coupled to the Vercel platform, so value drops sharply for services running elsewhere
- Aimed at application developers rather than platform teams, which shows in multi-team administration and policy
Best for: Product teams already deploying on Vercel who want to stop operating a proxy and whose requirements are one endpoint, failover and a spend ceiling.
Pricing: Usage-based on model tokens routed through the gateway within the Vercel platform.
OpenRouter

OpenRouter is worth considering if what you valued most about LiteLLM was breadth of model access rather than control. It is a hosted aggregator: one API, one key and one bill across a very large model catalogue and the providers serving them, with routing across those providers for availability and price and automatic failover between them. It is the only option here that competes with LiteLLM on coverage, and it does so by being the opposite of self-hosted.
Pros
- The only realistic managed answer to LiteLLM’s coverage advantage, including models you would otherwise need separate accounts for
- Routing across multiple providers serving the same model gives availability benefit a single-provider fallback cannot
- Billing aggregation removes real procurement work: one payment relationship instead of many
- Effectively zero adoption cost and zero operational load
Cons
- Hosted only, and every prompt traverses a third party, so it inverts the reason most teams self-hosted in the first place
- Not a governance product: RBAC, audit trails, budget enforcement hierarchies and prompt management are not what it does
- Adds a commercial intermediary between you and providers, affecting negotiated rates, enterprise terms and data-processing agreements
- Provider-level variation behind one model name means behaviour and latency can differ between requests in ways you do not control
Best for: Teams whose actual requirement was model breadth and consolidated billing, with governance and data control handled elsewhere or not yet required.
Pricing: Usage-based on tokens through a credit model, with the intermediary’s margin embedded in the per-model rate rather than charged as a platform fee.
TrueFoundry

TrueFoundry is the option when reasons one and two apply but you cannot give up the perimeter. Its control plane and data plane can run in your own Kubernetes cluster and VPC, so you get an enterprise governance surface without prompts leaving your network, and alongside the gateway it handles model deployment and serving — which matters if part of what LiteLLM was doing for you was fronting your own inference servers.
Pros
- Bring-your-own-cloud deployment satisfies data control and the enterprise governance requirement with one product
- Covers serving as well as gateway routing, the closest match if your LiteLLM deployment sat in front of your own models
- Enterprise-shaped on what blocks procurement — SSO, RBAC, audit trails, per-team budgets — rather than assembled
Cons
- Reduces operational load less than a fully managed gateway, because the deployment still lives in your cluster
- Substantially more platform than a team replacing a proxy usually wants, with a matching learning curve
- Commercial with a sales-led motion, and you take one vendor’s opinions across serving, routing and governance
- Loses LiteLLM’s config-as-code and no-exit-cost properties
Best for: Enterprises that need governance and support but cannot let payloads leave their own cloud account, particularly those already serving their own models.
Pricing: Commercial, sold on enterprise agreements rather than a public list price, with infrastructure on your own cloud account billed separately.
Envoy AI Gateway

Envoy AI Gateway answers a specific and quite common version of the complaint: you are happy self-hosting, you are not happy operating a Python service as a separate tier, and Envoy is already how traffic moves in your platform. It adds provider routing, upstream credential injection and token-aware rate limiting on top of Envoy Proxy and the Kubernetes Gateway API, configured with custom resources reviewed like the rest of your ingress.
Pros
- The data plane is Envoy, so the event-loop-blocking class of problem does not apply and streaming, pooling and timeouts behave as your platform team expects
- No new tier and no new database: LLM routing becomes configuration on infrastructure you already operate
- Kubernetes-native declarative config, so routing policy goes through the same review as the rest of your traffic rules
- Token-based rate limiting rather than request counting, the correct unit for model traffic
Cons
- No dashboard for spend, usage or prompt inspection — the admin UI and spend view you had become metrics, logs and your own assembly work
- Provider coverage is much narrower than LiteLLM’s, which is exactly the asset you would be trading away
- Assumes Kubernetes and real Envoy fluency; the CRD model is a steep fortnight for a team that just wanted a proxy
- Team hierarchies, budget workflows and prompt-level audit are not the shape of this project
Best for: Platform teams already running Envoy or Envoy Gateway who want to delete a tier rather than add a vendor, and who can live without a spend dashboard.
Pricing: Open source with no licence cost; the cost is cluster capacity plus platform engineering time to own the configuration and build the reporting.
Helicone
![]()
Helicone is worth a look if the part of LiteLLM you actually relied on was seeing what happened in each request. It approaches the path from the observability side: adopt with a base-URL change, get per-request logging, tracing, caching, rate limiting and cost attribution. It is self-hostable and open source, and it has announced it is joining Mintlify, a legitimate input into a multi-year decision about a component in your request path.
Pros
- Observability depth is the product, so prompt-level debugging and cost attribution are stronger than in a routing-first proxy
- Managed option removes operational load, self-hosted option keeps payloads inside your network — both reasons are addressable
- Very low adoption and removal cost, which makes it a safe intermediate step while you decide the bigger question
- Open source core, so the independence property you liked about LiteLLM is preserved
Cons
- Routing, fallback and multi-provider load balancing are not its centre of gravity, so it is not a like-for-like gateway replacement
- The full self-hosted deployment is several stateful services including a columnar analytics store — heavier than what you are leaving
- Governance depth — RBAC, budget enforcement, audit trails — is thinner than an enterprise gateway’s
- Its corporate home is changing, so apply the same roadmap scrutiny you would apply to any acquired component
Best for: Teams whose real requirement is per-request LLM visibility and prompt debugging, with routing simple enough to handle elsewhere.
Pricing: Open source and self-hostable with no licence cost, alongside a usage-based managed tier metered on logged requests.
How to choose
Name the reason first. Then read one row and ignore the rest of the table.
| Your reason for leaving | What actually qualifies | Where to start |
|---|---|---|
| Do not want to operate it | Fully managed, nothing stateful on your side | Portkey / Prisma AIRS, Cloudflare AI Gateway, Vercel AI Gateway, OpenRouter |
| Need enterprise governance out of the box | SSO, RBAC, audit export and budget enforcement as product | Portkey / Prisma AIRS, TrueFoundry, Kong AI Gateway |
| Want one gateway of record | LLM traffic as a route or plugin on your existing gateway | Kong AI Gateway, Envoy AI Gateway |
| Want the gateway at the edge | Proxy running near the caller, not in one cluster | Cloudflare AI Gateway, Vercel AI Gateway, Portkey / Prisma AIRS |
Then work through this before committing.
- Check whether reason one is really a staffing problem. If the deployment has no named owner, two replicas, a retention policy or a single retry layer, fix those first. Half the teams that want to migrate for operational reasons have never operated it properly, and every self-hosted alternative has the same requirements.
- Compare against LiteLLM’s enterprise tier, not its open-source build, if governance is your reason. That is the honest comparison and it sometimes ends the evaluation.
- Inventory the provider and model list you actually use, then check each candidate covers all of it including your own inference servers. This is the widest gap and the one that blocks a migration late.
- Test with your parameters, not the defaults. Reasoning controls, structured output, cache hints, tool definitions, multimodal payloads. OpenAI-compatibility is a floor; the parameters your product depends on are where candidates differ.
- Shadow real traffic to the top two and compare time-to-first-token per region rather than average latency. Then force a provider failure and watch the fallback, the retry count and what a streaming client receives.
- Cost the whole thing. Managed subscription plus token markup, against your current infrastructure plus the engineering hours you tracked. If you have not tracked the hours, that is the missing number, not the vendor quote.
The wider map is in the AI gateway hub, and the direct head-to-head is LiteLLM vs Portkey vs Kong AI Gateway. If routing sophistication specifically is what you want to improve, the model routing tools category is a more precise answer than a gateway swap.
Frequently asked questions
Is LiteLLM production-ready?
Yes, with the operational work any critical-path self-hosted service needs: two or more replicas, shared Redis for limits and cache, a Postgres with a retention policy, readiness probes that test dependencies, a termination grace period longer than your longest generation, and retries owned by exactly one layer. Teams that call it unreliable in production have usually skipped three of those.
Do I need to replace LiteLLM to get SSO and RBAC?
Not necessarily. Some of that sits in its paid enterprise tier, so the first comparison should be that tier against the managed platforms. You can also front the admin surface with an identity-aware proxy for SSO, though that covers gateway administration rather than giving you per-role policy inside the gateway, and it is internal glue nobody documents.
Will I lose provider coverage if I move to a managed gateway?
Almost certainly some, and this is the loss most likely to matter later. Managed gateways cover the major providers well and the long tail unevenly, and self-hosted inference servers vary from awkward to unsupported. Inventory the models you use today, add the ones you expect to evaluate next quarter, and check the list explicitly.
Is a Python proxy a problem at high concurrency?
It is a sizing and backpressure problem rather than a disqualifier. Watch CPU work on the event loop adding latency to unrelated requests, worker count against concurrent stream count, and whether the proxy buffers or applies backpressure with a slow consumer. Load-test on concurrent open streams with real payload sizes and you will know your ceiling. A gateway on Envoy or an edge runtime avoids that class of problem structurally.
Can I run LiteLLM behind another gateway instead of replacing it?
Yes, and it is a common and sensible end state. Your API gateway at the edge handles authentication, authorisation, audit and policy; LiteLLM behind it handles provider translation and model routing. You get one governance surface and keep the coverage advantage. The cost is two hops and two components, so it is worth it only when both jobs are genuinely needed — which, in an organisation with a real audit function and a wide model list, they usually are.
Related reading
- Best AI gateways — the full category map and where each type of gateway fits.
- LiteLLM vs Portkey vs Kong AI Gateway — the three architectural bets, compared directly.
- Best self-hosted LLM gateways — the operational checklist that fixes reason one without a migration.
- Portkey alternatives — the same decision from the other direction, segmented by reason.
- Best AI gateways for enterprise — the SSO, RBAC and audit requirements itemised.
- Best open source AI gateways — licence models and what each project keeps behind an enterprise tier.