The decision to self-host an LLM gateway usually takes about ten minutes. Someone points out that a managed gateway sees every prompt and every completion, legal agrees that is a problem, and the conclusion is “we’ll run it ourselves.” The container starts on the first try, the OpenAI-compatible endpoint answers, and everyone moves on.
What was actually decided in those ten minutes is that your team now owns a stateful multi-tier service sitting in front of every AI feature in the product. Not a proxy — a service with a database, a cache, a scaling story, a restart story, and a retention policy for the most sensitive text your users will ever type.
I have watched this pattern in other categories enough times to recognise it. Self-hosting is never about the software being hard to install. It is about the second and third components nobody put on the diagram, and about the fact that what you installed is now on the critical path of a product surface that did not exist eighteen months ago.
So this is not a ranking by feature count. It is what you operate after you say yes.
Key takeaways
- A self-hosted gateway is three components: a stateless proxy tier, a cache for rate limits and response caching, and a relational store for keys, budgets and request logs.
- That relational store is where prompt and completion payloads live. Turning on request logging is a data-retention decision, not a storage decision.
- The gateway is a hard dependency for every LLM feature. Two replicas, real health checks, and a termination grace period longer than your longest generation are not optional.
- Retries at the gateway on top of retries in the SDK turn one provider blip into a self-inflicted load spike. One layer owns retry policy; disable it everywhere else.
What you actually deploy
Every self-hosted LLM gateway converges on the same three-tier shape, whatever the project calls it.
A stateless proxy tier. A process that accepts an OpenAI-shaped request, resolves it to a provider and model, injects the real upstream credential, forwards the call and streams the response back. Holding no durable state, you scale it horizontally and treat instances as disposable. This tier is genuinely easy.
A cache, almost always Redis. Two jobs. Rate limits and budget counters need to be shared across replicas — a per-key limit enforced in each pod’s local memory is not a limit, it is your limit multiplied by your replica count, and it produces no error, just a bill. Response caching is the second, exact-match on a request hash or semantic against an embedding, and it needs somewhere that survives a restart.
A relational store, almost always Postgres. Virtual keys, team hierarchies, budget definitions, spend ledgers, model configuration, audit records, and — if you enable it — the full text of every request and response that passed through. This is the component people leave off the diagram, and it decides whether the deployment is still a success in a year.
Everything else the vendors advertise is a feature of one of those three tiers. Fallback chains, load balancing, token-based rate limits, guardrail hooks, spend dashboards: proxy logic plus one of the two stores.
The Postgres is a data-retention decision
Turn on request logging and your gateway’s database becomes a corpus of every prompt your users and systems have ever written, alongside every completion a model returned. Support transcripts, internal documents pasted into a chat box, code, customer records assembled by a RAG pipeline, whatever your agents put in a system prompt.
That is exactly the data you self-hosted to keep out of a vendor’s cloud. Self-hosting solved where it lives. It did not solve any of this:
- Retention. Prompt logs need a defined lifetime and something that enforces it. A partitioned table with a scheduled drop is the boring correct answer; a
DELETEjob over a growing unpartitioned table is how you meet autovacuum at an inconvenient moment. - Access. Read access to the gateway’s Postgres is a higher privilege than read access to most production tables, and it is rarely treated that way.
- Erasure requests. Prompt logs are in scope, so the logging schema needs an indexed tenant identifier, decided before you have a hundred million rows.
- Row size. Long-context requests are large, so full payload logging grows the table far faster than request count suggests, and it dominates your backup window before it dominates your disk.
The design that survives is two-tier: metadata and spend in Postgres indefinitely, full payloads either off, sampled, or written to object storage with a lifecycle rule. Decide before launch, because retrofitting means migrating your largest table.
Needs first-hand data: Enable full request logging on one replica for a representative day of your own traffic. Record bytes written per thousand requests, mean and p99 row size, and table growth per day. That number decides whether payload logging goes to Postgres, to object storage, or off.
If you already run a self-hosted logging and metrics stack, the lesson transfers — the self-hosted observability stacks comparison covers the same trap, where the storage tier rather than the collector decides the outcome. Send the gateway’s own metrics and traces there rather than standing up a second parallel stack, and keep its request logs as the sensitive tier that stays separate.
You are now in the request path
If the gateway is down, every LLM feature in your product is down. Chat, summarisation, classification, embeddings for search, whatever the agents do. One process class, total blast radius.
Health checks have to mean something. A liveness probe returning 200 because the HTTP server is listening will keep a pod in rotation while its connection pool is exhausted and every request fails. Readiness should test what the request path needs: Postgres reachable, Redis reachable, model config loaded. Liveness should be dumber than readiness, or a slow dependency triggers a restart storm.
Two replicas is the floor. One replica means every deploy, node drain and OOM kill is a customer-visible outage. The tier is stateless, so the only reason to run one is that someone still thinks of the gateway as a dev tool.
Rolling restarts truncate in-flight streams. This is the failure mode specific to LLM traffic and the one most likely to surprise a team with good Kubernetes practice. A streaming completion is a single long-lived HTTP response that may run for tens of seconds. On SIGTERM the default behaviour is to stop accepting connections and exit — and the client holding an open stream gets a truncated response, mid-sentence, with no error frame to distinguish it from a finished generation. Three things have to line up:
- A
preStophook and a termination grace period longer than your longest realistic generation. - A
PodDisruptionBudget, or the cluster autoscaler drains more of the tier at once than you intended. - Load balancer and ingress idle timeouts longer than your worst-case time-to-first-token. A proxy that times out an idle connection kills streams that were about to produce output, and the symptom looks like a model problem rather than a network policy.
Check your client too. The OpenAI-format stream ends with a defined terminator, and a client that treats “connection closed” as “generation complete” will silently show users half an answer.
Retry amplification is the failure you cause yourself. Provider SDKs retry certain error classes automatically. Your gateway also retries, and it has a fallback chain that tries a second provider, and your application has a retry wrapper somebody added during an incident. Multiply those and a thirty-second provider degradation becomes several times normal request volume, aimed at a provider already rate-limiting you, from a gateway whose connection pool is full of retries whose clients have given up.
The fix is architectural, not tuning. Pick one layer that owns retry policy — the gateway, because it is the only layer that can see the fallback options — and turn retries off everywhere else. Then add a circuit breaker per upstream, a cap on total attempts per request rather than per hop, and a deadline enforced end to end, so a request whose client disconnected stops consuming an upstream slot.
Needs first-hand data: Inject a 30-second failure on your primary provider with a fault-injection proxy and record total upstream request count versus client request count for that window. A ratio above two means retries are configured at more than one layer.
Air-gapped and egress-controlled deployments
There is one situation where a self-hosted gateway is not a preference but the only shape that works: when the network your models serve cannot reach the public internet. Defence, healthcare providers with strict residency rules, financial institutions whose egress allowlist is maintained by a security team that will not add a provider endpoint, and most government-adjacent deployments.
There the architecture is forced: a self-hosted inference server holding open-weight models on your own accelerators, a self-hosted gateway in front of it providing the OpenAI-compatible surface, key custody, per-team budgets and the audit trail, and nothing in the path that calls home. That is the cleanest deployment in this category, because the gateway’s job shrinks — no credential rotation across providers, no cross-provider fallback, no exchange-rate arithmetic on token pricing. What you give up is worth stating plainly rather than discovering in month three.
No hosted frontier models. You get what runs on your own hardware, and that ceiling moves every time a hosted model improves while your cluster stays the same.
No vendor dashboard, and no vendor. Spend visibility, usage analytics and the polished UI that sold the managed product are either absent or another service you now operate. Support is your team reading source code.
Distribution and upgrades become a process. Weights, container images and CVE patches all arrive by controlled transfer with an approval attached, so upgrading is a scheduled event rather than helm upgrade. While you are at it, check every component for telemetry pings and licence checks that assume connectivity — usually configurable, but you have to know to look.
If air-gapping is the driver rather than a preference, the constraint set is different enough that the compliance-focused gateway options are a more precise starting point.
The extra hop, honestly
A gateway adds a network hop, and whether that matters depends entirely on which call you mean.
On a multi-second generation the hop is noise. A chat completion streaming for four seconds does not care about single-digit milliseconds of proxy overhead. Anyone arguing a gateway is too slow for chat is arguing about a rounding error.
On an embedding call in a tight loop it is not noise. Embedding ten thousand chunks during an index build means ten thousand round trips, so the per-call overhead is multiplied by a number large enough to show up in wall-clock time. Same for a fast classifier called once per item in a stream job. Batch so the overhead amortises, or let the job talk to the provider directly and accept that it bypasses your spend accounting.
Three things decide whether the overhead is small or embarrassing, and all three are configuration. Connection reuse: a fresh TLS connection per request means a handshake on every call, so keepalive and a properly sized pool per provider is the difference between microseconds of CPU and a round trip. Streaming versus buffering: a proxy that reads the complete upstream response before forwarding turns a streaming API into a blocking one and makes time-to-first-token equal the full generation time — verify passthrough on day one, including through any ingress, mesh sidecar or WAF in the path, each of which can buffer too. Placement: an extra hop within a region is cheap, across regions it is a real round trip on every call, and one gateway serving applications in three clusters is a common accident.
Semantic caching deserves a warning, because it is sold as a latency win and is not always one. It embeds the incoming request to find a similar previous one, so a cache miss costs an embedding call plus a vector search before the real request starts. At a high hit rate that is a large win; at a low hit rate it is pure added latency and cost. Understand the semantic caching options before enabling it by default.
Needs first-hand data: Measure p50 and p99 added latency for three call shapes through your gateway versus direct to provider — a streaming chat completion, a single embedding call, and a batch of 100 embeddings — reporting time-to-first-token separately for the streaming case. Those six numbers settle the “is the gateway slow” argument permanently.
The OpenAI API format: the standard underneath, not a product
Every gateway below advertises an OpenAI-compatible endpoint, and it is worth being precise, because that is not a product you choose or pay for.
The OpenAI request and response shape — a messages array, tool definitions, an SSE stream of deltas at a /v1/chat/completions-style path — became the de facto wire format for talking to a language model. Provider SDKs speak it, inference servers implement it, and every gateway here presents it as its front door. That is why a self-hosted gateway can be dropped in front of an application by changing a base URL rather than rewriting call sites.
What it gives you
- One client library and one request shape across hosted providers and your own inference servers
- A drop-in insertion point: pointing an existing application at a gateway is a base-URL change
- Portability at the application layer, so replacing the gateway later does not touch your call sites
- A common vocabulary for tools, streaming and structured output that most of the ecosystem now assumes
What it does not do
- Compatibility is partial. Provider-specific parameters, reasoning controls, cache hints, safety settings and multimodal payloads differ, and a gateway either drops them, passes them through, or translates them with judgment calls you inherit
- Error semantics are not standardised. Rate-limit signalling, retry-after headers and content-filter responses vary, and how your gateway normalises them determines whether your retry logic is correct
- It says nothing about authentication, quotas, audit or multi-tenancy — which is the entire reason a gateway exists
Treat OpenAI-compatibility as a floor that gets you connected, then test the parameters your product depends on against each provider through your specific gateway. The OpenAI-compatible proxy landscape goes deeper on where the abstraction leaks.
LiteLLM

LiteLLM is the default answer for self-hosting, and deservedly so. It began as a Python library normalising a very long list of provider APIs to the OpenAI call shape, then grew a proxy server around that translation layer. Self-hosted is its native mode: run the proxy container, point it at a Postgres for keys, budgets and spend and a Redis for shared limits and caching, and declare models in YAML. Provider coverage is its main asset, and it treats your own inference server as just another entry.
Pros
- The broadest provider and model coverage in the category, which is what you want in front of a fast-moving model market
- Genuinely self-host-first: virtual keys, team budgets, spend tracking and the admin UI are in the open-source proxy
- Configuration is YAML in version control, so model lists and fallback chains are reviewable artifacts rather than console state
- Self-hosted inference servers are first-class, which is what makes the air-gapped shape straightforward
Cons
- Python in the hot path means the proxy tier needs deliberate worker and concurrency sizing under high request rates with many concurrent streams — the LiteLLM alternatives piece covers what that involves
- You own three components, and the Postgres full of prompt payloads is your retention problem
- Large, fast-moving surface area, so pinning versions and reading release notes is ongoing work
- Some enterprise governance sits behind a paid tier, so “open source” is not the whole feature list
Best for: Teams that want maximum provider coverage and a self-hosted control plane with keys and budgets included, and who have someone willing to own three components as a production service.
Pricing: Open source with no licence cost for the core proxy plus a paid enterprise tier for advanced governance and support; self-hosting moves the cost into infrastructure and operator time.
Envoy AI Gateway

Envoy AI Gateway is for teams whose answer to “what proxies our traffic” is already Envoy. It builds LLM-specific behaviour — provider routing, upstream credential injection, token-aware rate limiting — on top of Envoy Proxy and the Kubernetes Gateway API, configured with custom resources rather than a bespoke file. The data plane is the same proxy already handling your service-to-service traffic, so streaming and connection behaviour are a known quantity rather than something new to evaluate.
Pros
- The data plane is Envoy, so streaming, connection pooling, timeouts and observability behave the way your platform team expects
- Kubernetes-native declarative configuration, so LLM routing is reviewed like the rest of your ingress
- Token-based rate limiting rather than request counting, which is the correct unit for LLM traffic
- No separate vendor control plane to trust or operate; it is infrastructure you already run, extended
Cons
- No product-grade dashboard for spend, usage or prompt inspection — you assemble that from metrics and logs
- Assumes Kubernetes and real Envoy fluency; the CRD model is a steep first fortnight for a team that just wanted a proxy
- Team hierarchies, budget approval flows and prompt-level audit UI are not the shape of this project
Best for: Platform teams already running Envoy or Envoy Gateway in Kubernetes who want LLM routing as another data-plane concern rather than a new service with its own database.
Pricing: Open source with no licence cost; the cost is cluster capacity plus the platform engineering time to own the configuration.
Apache APISIX

APISIX is a general-purpose API gateway from the Apache Software Foundation, built on Nginx and OpenResty with etcd as its configuration store, and it has grown a family of AI plugins — provider proxying, multi-provider routing, prompt guarding, prompt templates and token-aware rate limiting. The point is that LLM traffic is handled by plugins on the gateway that already handles your REST APIs, so you extend an existing control plane rather than introducing a second one.
Pros
- One gateway, one config store and one set of operational habits for both REST and LLM traffic
- Foundation-governed rather than single-vendor, which matters for teams with concerns about a project’s direction
- Dynamic configuration through etcd applies plugin and route changes without restarts, which is what you want in a component holding long-lived streams
- Mature Nginx-based data plane with well-understood streaming and connection behaviour
Cons
- The AI plugin family is newer than the gateway, so spend ledgers, per-team budgets and prompt-level analytics trail purpose-built LLM gateways
- etcd is another stateful component to operate, back up and upgrade correctly
- Custom logic means Lua or a plugin runner, a narrower skill pool than Python or Go on most teams
Best for: Teams already running APISIX for their APIs who want LLM routing and token-based limits as additional plugins rather than a new tier.
Pricing: Open source with no licence cost, with commercial support and managed offerings available in its ecosystem; self-hosted cost is gateway nodes plus etcd.
Kong AI Gateway

Kong AI Gateway is Kong Gateway plus a set of AI plugins, making the same bet as APISIX with a stronger enterprise posture: LLM traffic is another protocol governed by the API gateway you may already run. Its product framing extends past model calls to MCP and agent-to-agent traffic through the same gateway, which is a meaningful position if you can see agent traffic becoming a governance problem. Self-hosting is first-class, in DB-less declarative mode or backed by Postgres.
Pros
- One gateway of record for REST, LLM, MCP and agent-to-agent traffic, with one policy layer and a team that already knows the tooling
- The AI plugins cover the practical needs — provider proxying, load balancing across models, prompt guarding, semantic caching, request and response transformation — as configuration rather than code
- DB-less declarative deployment removes a stateful component if you do not need the full feature set
- Existing authentication, authorisation, logging and rate-limiting plugins apply to LLM routes unchanged
Cons
- The open-source and enterprise plugin split is the first thing to check; several advanced AI capabilities are enterprise-tier, so an open-source evaluation may not match what you ship
- If you do not already run Kong, adopting it for LLM traffic alone is a lot of gateway for a narrow purpose
- LLM-specific analytics and spend attribution are less developed than in gateways whose entire product is model traffic
Best for: Platform teams already running Kong who want one control plane and one policy layer covering REST, LLM, MCP and agent traffic rather than a second LLM-only hop.
Pricing: Open-source gateway with no licence cost plus an enterprise subscription for the advanced plugin set and managed control plane; which AI plugins fall on which side of that line is the question that decides your cost.
TrueFoundry

TrueFoundry is a platform rather than a proxy, and its relevance here is the deployment model: a control plane and data plane that can run inside your own Kubernetes cluster and VPC, so payloads never leave your network. Alongside the gateway it does model deployment and serving, making it the closest thing here to a single answer for “run models and govern access to them, all inside our perimeter.”
Pros
- Designed for bring-your-own-cloud deployment, so enterprise governance and data residency are satisfied by the same product
- Covers both serving your own models and gateway-level routing, budgets and access control, reducing vendor count inside the perimeter
- Enterprise-shaped on what blocks procurement: SSO, role-based access, audit trails, per-team budgets
Cons
- Considerably more platform than a team that wanted a proxy is prepared to adopt, with a matching learning curve
- Commercial with a sales-led motion, so evaluation is a conversation rather than a
docker run - One vendor’s opinions across serving, routing and governance rather than composable pieces you can replace individually
Best for: Enterprises that need both model serving and gateway governance inside their own VPC, with SSO and audit as hard requirements.
Pricing: Commercial platform on enterprise agreements rather than a public self-serve list price, with the deployment running in your own cloud account so infrastructure is billed separately.
Helicone
![]()
Helicone approaches the same request path from the observability side. Its original shape was a proxy you adopt with a base-URL change, which then gives you logging, tracing, caching and rate limiting over your LLM calls, and it is self-hostable. Helicone has announced it is joining Mintlify, worth factoring into a multi-year self-hosting decision as you would any change of ownership. Note that the full self-hosted deployment is heavier than a single container, because the analytics side needs a columnar store behind it.
Pros
- Very low adoption cost: a base-URL change gets request-level visibility without restructuring how your code calls models
- Observability depth — per-request traces, prompt inspection, cost attribution — is the product rather than a side feature
- Self-hostable, so the prompt payloads it is designed to capture stay inside your network
Cons
- The complete self-hosted stack is multiple stateful services including a columnar analytics store, heavier than proxy plus Postgres plus Redis
- Routing, fallback and multi-provider load balancing are not its centre of gravity, so as a pure gateway it does less
- Its corporate home is changing, a legitimate roadmap consideration for infrastructure you intend to run for years
Best for: Teams whose primary requirement is per-request LLM visibility and prompt-level debugging inside their own infrastructure, with routing secondary.
Pricing: Open source and self-hostable with no licence cost, alongside a usage-based managed offering metered on logged requests.
vLLM

vLLM is not a gateway and should not be evaluated as one. It is an inference server: it loads open-weight models onto your accelerators, uses paged attention and continuous batching to keep utilisation high across concurrent requests, and exposes an OpenAI-compatible endpoint. It belongs here because in a self-hosted deployment vLLM is what sits behind the gateway — the gateway does keys, budgets, routing and audit, vLLM does the tokens. Conflating the two is how teams end up with no access control in front of a GPU cluster.
Pros
- The performance-oriented default for self-hosted open-weight inference, keeping expensive accelerators busy
- OpenAI-compatible server surface, so it plugs into any gateway here as just another provider entry
- Runs entirely inside your network, which is what makes the air-gapped shape possible at all
Cons
- No multi-tenancy, virtual keys, budgets, spend attribution or audit trail — everything a gateway provides is absent by design
- Operating it means operating GPU infrastructure: capacity planning, model load times, per-model memory sizing, queueing under load
- It serves the models it has loaded, so cross-provider routing and fallback remain the gateway’s job
Best for: Any self-hosted deployment that needs open-weight models on its own hardware, behind a gateway that provides the access control it deliberately lacks.
Pricing: Open source with no licence cost; the cost is the accelerators it runs on and the engineering time to size and operate them.
How to choose
Start from what you already operate. The gateway that fits is almost always the one that adds the fewest new components.
| Option | What it is | New stateful components | Picks itself when |
|---|---|---|---|
| LiteLLM | LLM-first proxy with keys, budgets and spend built in | Postgres and Redis | You want maximum provider coverage and a control plane out of the box |
| Envoy AI Gateway | LLM routing on the Envoy data plane, via Gateway API | None beyond your cluster | Envoy is already how traffic moves in your platform |
| Apache APISIX | General-purpose gateway with an AI plugin family | etcd | You already run APISIX for REST APIs |
| Kong AI Gateway | Gateway of record for REST, LLM, MCP and A2A traffic | Postgres, or none in DB-less mode | Kong is already your policy layer |
| TrueFoundry | BYOC platform covering serving and gateway governance | Managed by the platform in your cluster | You need serving plus enterprise governance in your own VPC |
| Helicone | Observability-first proxy, self-hostable | Postgres plus a columnar analytics store | Per-request prompt visibility is the real requirement |
| vLLM (inference server, not a gateway) | Serves open-weight models on your accelerators | None; it is the upstream | You are running your own models behind one of the above |
Then run this sequence, which is mostly not about the gateway.
- Write down which of the three tiers you are prepared to operate. If the answer is “the proxy only”, you want the Envoy or APISIX shape — or a managed gateway and a different article.
- Decide the payload logging policy before you deploy. On, off, or sampled to object storage with a lifecycle rule. This is the expensive decision to reverse.
- Deploy two replicas from the start, then kill a pod mid-stream and watch your client. If the UI shows a truncated answer as complete, fix that before the grace period and the disruption budget.
- Audit retry configuration at every layer and disable all but one, then inject a provider failure and count upstream requests.
- Measure the hop for your three worst call shapes and decide which paths may bypass the gateway. Only then compare features — by now you will have eliminated most of the list on operational grounds, which is the right order.
The broader menu, including the managed options this article skips, is in the AI gateway hub, and the open-source gateway roundup covers licence and community questions in more depth.
Frequently asked questions
Do I really need Redis, or can I skip it?
You can skip it with one replica, which means you should not skip it. Rate limits and budget counters held in one process’s memory are enforced per process, so the moment you scale to two pods every limit silently doubles. A shared store is what makes horizontal scaling correct rather than merely possible.
Should the gateway log full prompts and completions?
Default to no, then enable it deliberately where it earns its place. Debugging and evaluation genuinely benefit from full payloads, so the useful middle ground is sampling, or full logging for a short retention window with metadata and spend kept indefinitely. What you should not do is turn it on globally with no retention plan and discover the corpus during an audit.
Does self-hosting the gateway mean self-hosting the models?
No, and most deployments do not. A self-hosted gateway calling hosted provider APIs is the common case: key custody, spend accounting and audit stay inside your perimeter while prompts still travel to the provider. Self-hosting the models too is what a strict-egress or air-gapped environment requires, and that is a much larger commitment.
Related reading
- Best AI gateways — the full category map including the managed options this article leaves out.
- Best open source AI gateways — licence models and what “open source” excludes in each project.
- LiteLLM vs Portkey vs Kong AI Gateway — the three architectural bets, compared directly.
- LiteLLM alternatives — what to do when the operational load is the thing you want to shed.
- Best self-hosted observability stacks — the same storage-tier trap, in the category you will use to monitor the gateway.
- Best AI gateways for compliance — when air-gapping and data residency drive the whole decision.