Buyer’s Guide

Best Open Source AI Gateways

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • ai-gateway
  • open-source
  • self-hosted
  • llm

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

“Open source AI gateway” is doing a lot of work in most of these comparisons. Almost every product in this category has a repository you can clone, and almost every one of them also has a page listing the features that are not in it. The interesting question is never whether the project is open. It is where the line runs, and whether the features you specifically need sit above or below it.

The line is remarkably consistent across the category, which is what makes it predictable. Routing, provider translation, caching and basic key management are open. SSO, fine-grained RBAC, audit logging, hard budget enforcement and high-availability clustering are frequently not. That is not a coincidence or a criticism — those five are exactly the features that turn a proxy into something an enterprise will sign off, which makes them the natural place to put a price.

So this article is two procedures and six products. The first procedure establishes where a project’s paid line sits before you build against it. The second reads the licence, because “open source” covers three legally different situations that behave very differently the day you patch the code or run it as a service for someone else.

Key takeaways

  • Five features are the classic paid line here: SSO, fine-grained RBAC, audit logging, budget enforcement and HA clustering. Check each one specifically rather than trusting the word “open”.
  • Licence shape matters more than the label. Permissive, copyleft and source-available-with-restrictions behave differently when you patch the code or offer it as a service.
  • Foundation-governed projects and single-vendor projects are different bets. When a single vendor is acquired the repository usually survives and the roadmap changes.
  • Free of licence cost is not free. The bill is a stateful service in your request path, its upgrades, and someone on call for it.

Read the open-core boundary before you commit

Do this before you write any integration code. It takes an hour and it is the difference between a self-hosted gateway and a self-hosted gateway plus a procurement cycle you did not plan.

Find the feature matrix, not the README. The README describes the project. The pricing or enterprise page describes the boundary. If a vendor has an open-source project and a commercial offering, there is a comparison table somewhere; that table is the actual specification of what you get for nothing. If you cannot find one, treat that as a finding rather than an absence — a boundary that is not documented is a boundary that can move.

Then check these five, individually. Everything else varies by product; these five are where the line usually sits.

  • SSO. Whether the self-hosted build authenticates administrators against your identity provider, or only supports local accounts and static keys. Local accounts are workable for a platform team of three and fail an access review immediately.
  • Fine-grained RBAC. Whether roles can distinguish who changes routing, who reads prompt contents in logs and who raises a budget, or whether there is one admin role that does everything. One role is not RBAC.
  • Audit logging. Whether administrative actions — a routing change, a budget increase, a key issued — produce a tamper-evident, exportable record. Request logs are not audit logs; they record what the system did, not who told it to.
  • Budget enforcement. Whether the open build actually blocks a call at the limit or only records spend. Enforcement is often the paid half of a feature whose reporting half is open, and the distinction is easy to miss in a demo.
  • High-availability clustering. Whether you can run multiple instances that share rate-limit counters, budget state and cache correctly, or whether the open build is effectively single-instance. This one decides whether the gateway can be in a production request path at all.

Check how the boundary is enforced in the code. Two patterns, differing in how safely you operate. Some projects keep enterprise features in the same repository behind a licence check: the code is visible, the upgrade path is a key, and you must not use those paths without a licence. Others keep them in a separate distribution entirely. The first is more convenient and puts an obligation on you to know which code paths you may execute.

Check who can change the boundary. A single company owning the copyright can relicense or move a feature behind the paid line in a future version. A foundation-governed project with contributions from multiple companies cannot do that unilaterally. This is the most important structural difference here, and it is invisible on a feature matrix.

Then check the boring things that decide operational cost. Release cadence and whether a supported long-term version exists. Whether upgrades have ever required a migration. Whether images are published for the architectures you run. Whether there is a security disclosure process — a gateway in your request path with no published patch process is a risk you accept on behalf of every service behind it.

Needs first-hand data: For each candidate, deploy the open build with two replicas behind a load balancer, set a shared rate limit and a budget, then drive concurrent traffic through both. Record whether the limits are enforced globally or per instance. That single test separates the projects that can front production from the ones that cannot.

Licence shape matters more than the word “open”

Three legally distinct situations get called open source in this category, and they diverge exactly where it matters: when you modify the code, and when you run it as a service.

Permissive. You can use, modify, embed and redistribute with attribution, including inside a commercial product, without publishing your changes. For a gateway this is the least constrained option: patch it, keep the patch private, ship it in your own platform. The obligation is essentially attribution and notice.

Copyleft. You keep the same freedoms, and derivative works you distribute must carry the same licence. The version that matters for a gateway is the network-service form of copyleft, which treats offering the software to users over a network as distribution. If you patch such a gateway and expose it to your own customers as part of a product, the obligation can reach your modifications. For purely internal use the practical impact is usually small; for a product you sell, this is a conversation with a lawyer rather than an engineer.

Source-available with restrictions. You can read the code, run it and modify it, but the licence forbids some class of use — most commonly offering it as a competing managed service, sometimes with a delayed conversion to a permissive licence after a fixed period. This is legitimate and it is not open source in the sense most engineers assume. It is usually fine if you are an end user, and it is disqualifying if your product is a platform that would resell the capability.

Three practical consequences, whichever you land on.

Patching. Under a permissive licence a private fork is entirely yours. Under network copyleft, a fork you expose to customers may carry publication obligations. Under source-available, a fork is fine but the use restriction still applies. In every case a private fork means you own the merge conflict with every upstream release, which is a cost independent of the licence.

Running it as a service. If you are building a platform other companies pay for and the gateway is part of what they pay for, read the licence yourself rather than a summary. This is the one scenario where the three shapes give different answers, and exactly the scenario a platform team is usually in.

Contributing back. If you need a fix upstream, the contributor agreement matters. Some projects ask contributors to assign broad rights to one company — normal for single-vendor open source, and worth knowing before your engineers spend a week on a patch.

Single-vendor open source is a maintenance bet

The most useful thing to understand about this category is that most of its open-source projects have exactly one company behind them, and the pattern when that company changes hands is now well documented.

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks. Helicone has announced it is joining Mintlify. In adjacent categories the same thing has happened repeatedly: Promptfoo is now part of OpenAI and still ships its open-source CLI, Traceloop is joining ServiceNow with OpenLLMetry as its open instrumentation layer, Arize has announced a new chapter with Dynatrace while Phoenix continues as its open-source tracing and eval project, Galileo is now part of Cisco, and OpenMeter is now OpenMeter by Kong.

The pattern across those cases is consistent: the repository survives, and the roadmap changes. Nobody deletes the project — that would be reputationally expensive and pointless. What changes is priority. Features serving the acquirer’s platform get built; features serving the long tail of self-hosting users compete with an enterprise integration roadmap. Release cadence usually slows before it recovers. The community edition keeps working and stops being where the interesting work happens.

That is survivable for most teams and it is a real input to a decision. Two things follow.

Judge the bet on governance, not enthusiasm. A project with contributions from several companies under a foundation cannot be repointed by one acquisition. A project with one corporate contributor can. Neither is wrong; know which you signed up for.

Keep the integration shallow either way. If your services talk to the gateway over the OpenAI format, replacing it is a base-URL change and a policy port. If your code imports the gateway’s SDK and depends on its config semantics, replacing it is a project. That difference is within your control and it is the cheapest insurance available.

Needs first-hand data: For each candidate, count distinct organisations among the top twenty contributors over the last two release cycles, and record time-to-first-response on the last ten security-labelled issues. Those two numbers predict maintenance risk better than any star count or launch announcement.

LiteLLM

LiteLLM homepage

LiteLLM is the most widely used open project here and the clearest example of the open-core pattern. The core is a Python library plus a proxy speaking the OpenAI format in front of a very long provider list, with virtual keys, spend logging and per-key budgets. Around it sits a commercial enterprise tier, and the features an organisation needs at scale are the ones most likely to sit there — which makes it the strongest case for running the boundary procedure first.

Pros

  • Widest provider coverage of any self-hosted option, local runtimes and private endpoints included
  • Library and proxy are one project, so you can adopt in-process and move to a hop without changing call sites
  • Virtual keys with spend logs give per-team attribution with no vendor in the path
  • Large user base, so failure modes you hit are usually already discussed somewhere

Cons

  • Single-vendor open core, so the paid line can move; check where SSO, RBAC, audit logging and enforcement sit today
  • The proxy is a stateful service needing a database, and running it highly available is more work than the quickstart implies
  • Breadth brings uneven quality: heavily used providers are solid, rare ones carry less-exercised code

Best for: Teams that want the broadest self-hosted provider translation layer and will verify the open-core boundary against their own requirements first.

Pricing: Open source with no licence cost for the core plus a paid enterprise tier for organisation-level features and support; the rest is infrastructure and your own time. LiteLLM alternatives covers the exits if the boundary does not fit.

Envoy AI Gateway

Envoy AI Gateway homepage

Envoy AI Gateway changes the governance answer rather than the feature list. It applies provider credentials, model routing and token-aware rate limiting as a layer over an Envoy data plane, built in the Envoy ecosystem under open governance rather than by one company. For the maintenance bet above, that is the structural difference: no single vendor whose acquisition repoints the roadmap, and no commercial tier holding the features an enterprise needs.

Pros

  • Multi-vendor governance rather than one company’s open core, removing both paid-line and acquisition risk
  • The data plane is Envoy, already understood and already in the request path in most Kubernetes estates
  • Kubernetes-native config fits GitOps and policy review, so change management is the process you have
  • Token-aware rate limiting is first-class rather than approximated from request counts

Cons

  • Assumes real Envoy and Kubernetes competence; the hardest option here without a platform team
  • No commercial tier means no vendor to buy support from, which some organisations cannot accommodate
  • Prompt management, dashboards and evaluation are out of scope by design and assembled separately

Best for: Kubernetes platform teams with real Envoy expertise who want LLM policy under governance no single vendor controls.

Pricing: Open source with no licence cost and no paid tier; the entire cost is infrastructure plus the engineers who operate it.

Apache APISIX

Apache APISIX homepage

APISIX is a general-purpose API gateway developed under Apache Software Foundation governance, with plugins that extend it to LLM traffic. It belongs here for the same reason as Envoy: foundation governance means the project is not one company’s asset, so the open-core question does not arise in the same form. Be careful for one reason — its LLM-specific plugins are much newer than its core, so verify the behaviour you need rather than the plugin’s existence.

Pros

  • Foundation governance with multiple contributing organisations, so no single vendor can move the paid line
  • Mature, well-exercised gateway core: routing, authentication, rate limiting and observability at volume
  • Dynamic configuration without restarts, which makes routing and policy changes operationally cheap
  • One gateway for REST and LLM traffic, so you are not adding a second control plane

Cons

  • The LLM plugin surface is newer than the core, so token accounting, streaming and provider coverage need verifying
  • Plugin development is a specific skill, and extending it is a bigger commitment than configuring a purpose-built gateway
  • No prompt lifecycle, evaluation or per-tenant cost tooling; a traffic control plane only

Best for: Platform teams that want a foundation-governed gateway for both REST and LLM traffic and will test the AI plugins against their own workload.

Pricing: Open source with no licence cost; commercial support and managed offerings exist from vendors in the ecosystem, and the project itself is free of licence obligation beyond notice.

Kong

Kong homepage

Kong is the open-core API gateway most enterprises here already run, with AI plugins extending it to model traffic — and its product now covers LLM, MCP and agent-to-agent traffic through the same gateway. Its shape is familiar: a capable open core with paid enterprise tiers, backed by a single company. If you already run it, the AI plugins are the cheapest governed path available, because the identity model, audit pipeline and change process are in place.

Pros

  • Where Kong is already deployed, adding LLM governance is configuration rather than a new system
  • One policy plane across REST, LLM, MCP and agent traffic under a single identity model
  • Mature operations: highly available deployment, multi-workspace tenancy and upgrade paths exercised at scale
  • Large plugin ecosystem, so adjacent requirements often have an existing answer

Cons

  • Single-vendor open core, so enterprise-shaped features sit above a line that vendor can move
  • Heavy to adopt purely for LLM traffic: an API gateway platform taken on to get a gateway feature
  • Configuration-first, so application teams will want something in front of it

Best for: Platform teams already running Kong who want LLM, MCP and agent traffic governed by the control plane they already operate.

Pricing: Open-source core with paid enterprise tiers priced by deployment scale and support; the AI plugins follow the same split.

Helicone

Helicone homepage

Helicone’s value is observability rather than routing: point a base URL at it and every request, response, token count and latency is recorded, with caching and rate limiting at the same hop. The self-hostable build is why it appears here, because it answers the objection that prompts must not leave your network. It has announced it is joining Mintlify, which makes it a live example of the maintenance bet rather than a hypothetical one.

Pros

  • Fastest path from zero to per-request LLM logs, with no instrumentation to write anywhere
  • Self-hostable, so prompt content stays inside your own infrastructure
  • Header-based session and user attribution makes per-customer cost answerable without a data project

Cons

  • Now inside a documentation company’s portfolio, so the self-hosted roadmap is a genuine unknown for a long dependency
  • Proxy-only visibility: agent steps, retrieval and tool executions never reach it
  • Routing, fallback and hard budget enforcement are lighter than in policy-designed gateways

Best for: Teams that need self-hosted visibility into LLM traffic quickly and can accept roadmap uncertainty on the open build.

Pricing: Open source and self-hostable at infrastructure cost, with a managed tier metered on logged requests and retention. Deeper tracing options are in LLM observability tools.

GPTCache

GPTCache homepage

GPTCache is not a gateway and should not be evaluated as one — it is an open-source semantic caching library. You embed it in your application or run it alongside, and it answers repeated or near-identical requests from a cache using embedding similarity rather than exact byte matching. It appears here because caching is one of the capabilities people expect from a gateway, and if caching is the only capability you need, a library is a smaller commitment than a proxy. Check the project’s recent commit activity before you build on it, as you would for any library in your request path.

Pros

  • Solves one problem without introducing a network hop, a service to operate or a new failure mode
  • Semantic matching catches paraphrased repeat questions that exact-match caching in a gateway misses
  • Pluggable storage and embedding backends, so it fits the infrastructure you already run
  • Library scope means the decision is reversible: removing it is deleting a dependency, not migrating a platform

Cons

  • Not a gateway: no routing, no fallback, no key custody, no budgets, no audit trail
  • Semantic caching introduces a correctness risk that belongs to your product — a similar question is not the same question, and the similarity threshold is a product decision with user-visible consequences
  • In-process caching does not share hits across services, so each service warms its own cache unless you back it with shared storage

Best for: Teams whose only unmet need is caching repeated model calls, who would rather add a library than operate a proxy.

Pricing: Open source with no licence cost; you pay for the cache storage and embedding calls it uses. The wider category is covered in semantic caching tools.

How to choose

Run the boundary procedure first, then this.

Start with governance, because it is the thing you cannot change later. If your requirement is that no vendor’s commercial strategy can affect your gateway, that eliminates the open-core products and leaves the foundation-governed ones. If you are comfortable with a single-vendor project and want breadth, the calculation is different.

Then check the five paid-line features against your actual requirements. Not against a future state — against what you need in the next two quarters. A team of four with no access review does not need SSO today, and should still know whether it is available without a contract, because that answer decides whether adoption ends in a procurement cycle.

Then test high availability specifically. Two replicas, one shared rate limit, concurrent traffic. This is the test that most often reveals that the open build is single-instance in practice, and it is the one that decides whether the gateway can carry production traffic.

Then read the licence, once, properly. Especially if you will patch the code or expose the capability to paying customers. Ten minutes of reading is cheaper than the alternative.

ProjectGovernanceWhere the paid line usually bitesPicks itself when
LiteLLMSingle vendor, open coreOrganisation-level controls and supportYou need the widest provider coverage self-hosted
Envoy AI GatewayMulti-vendor, open governanceNo paid tier to biteYou run Envoy and want no vendor dependency
Apache APISIXApache Software FoundationNo paid tier in the project itselfYou want one foundation-governed gateway for REST and LLM
KongSingle vendor, open coreEnterprise governance featuresKong already fronts your traffic
HeliconeSingle vendor, now inside an acquirerManaged features and scaleYou need self-hosted visibility this week
GPTCache (library, not a gateway)Open-source libraryNot applicableCaching is the only capability you need

The architectural background to all of this is in the AI gateways hub. If the deciding factor is operating the thing rather than licensing it, self-hosted LLM gateways quantifies that load, and self-hosted observability stacks is the closest analogue for what running your own always-on infrastructure costs in practice. If enterprise controls are the constraint, AI gateways for enterprise is the version of this shortlist that starts from procurement.

Frequently asked questions

Is an open source AI gateway cheaper than a managed one?

Cheaper in licence cost, not automatically cheaper overall. You are taking on a stateful service in the path of your most important calls, plus its database, upgrades, capacity planning and an on-call rotation. At meaningful volume the arithmetic usually favours self-hosting; at low volume it usually does not, and the crossover is much higher than teams expect because the dominant term is engineer time.

Which features are most often behind the paid line?

SSO, fine-grained RBAC, audit logging, hard budget enforcement and high-availability clustering. Those five recur across the category because they are what an enterprise buyer requires, which makes them the rational place to charge. Verify each against the specific product and version you intend to run rather than assuming from the category pattern.

Does an acquisition mean the open source project will be abandoned?

Usually not. Across the acquisitions in this space the repositories have continued to exist and to work. What changes is the roadmap and often the release cadence, as the project’s priorities align with the acquirer’s platform. Plan for a slower-moving dependency rather than a dead one, and keep your integration shallow enough to leave.

Can I run a gateway with a permissive licence inside a product I sell?

Under a permissive licence, generally yes, subject to attribution and notice. Under a network-copyleft licence, exposing a modified version to your users can carry publication obligations. Under a source-available licence, offering it as a competing service is typically the restricted case. This is the one question in this article where you should read the actual text rather than a summary of it.