Buyer’s Guide

Best API Gateways for Kubernetes

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • api
  • gateway
  • kubernetes
  • ingress
  • infrastructure

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

Every Kubernetes cluster already has something doing this job. Usually it is an ingress controller somebody installed in week one, configured with a handful of annotations, which has since accumulated rewrite rules, timeout overrides and a CORS header that nobody remembers adding. It works. Then a team needs per-consumer rate limits, or you need to split traffic for a canary, or an auditor asks who is allowed to change routing, and the annotation soup stops being sufficient.

That is the real starting point for this decision, and it is why the question is rarely “which gateway” and usually “do we extend what we have, replace it, or admit that the service mesh we already installed does half of this”.

Three things make the Kubernetes version of this decision different from the general one. The configuration surface is Kubernetes objects rather than a config file, so the ergonomics of expressing per-route policy in YAML matter enormously. There is a real standard now in Gateway API, which changes the portability calculation. And most of the credible options are Envoy underneath, which means they differ in control plane and operational model far more than in data plane behaviour.

Key takeaways

  • Gateway API is a specification, not an implementation. Adopting it changes who owns which part of routing configuration, which is its main benefit; it does not by itself give you policy features.
  • Ingress annotations are the thing you are escaping. They are controller-specific, untyped, invisible to RBAC at a useful granularity, and the source of most routing incidents in mature clusters.
  • A service mesh already terminates external traffic in most installs. Running a mesh gateway plus a separate ingress controller is common and frequently unnecessary duplication.
  • Most options here are Envoy with a different control plane attached. Choose on the control plane’s operational complexity and config ergonomics, because the proxy is largely the same.

Gateway API: a specification, not an implementation

Gateway API is the Kubernetes project’s successor to Ingress. It is a set of custom resource definitions and a conformance suite, not a piece of software you run. Several of the products below implement it, with differing completeness, and you still have to pick one of them.

Its central idea is role separation. Ingress collapsed everything into one object that a developer edited and a cluster admin hoped was correct. Gateway API splits it into three: GatewayClass is infrastructure, owned by the platform team and pointing at a controller implementation. Gateway is a listener with ports, protocols, TLS configuration and rules about which namespaces may attach to it, also platform-owned. HTTPRoute (and GRPCRoute, TCPRoute and friends) is the routing rule, owned by the application team in their own namespace, attached to a Gateway that permits it.

That split is the actual product. It means a team can add a route without being able to change TLS termination or bind a new port, and it means the permission boundary is expressed in Kubernetes RBAC on distinct resource types rather than in a review convention.

What it gives you

  • Typed, first-class fields for header matching, traffic splitting by weight, request mirroring, timeouts and redirects, all of which were annotations before
  • A real permission boundary: route authors cannot change listener or TLS configuration, and Gateways control which namespaces may attach routes
  • Cross-namespace routing with an explicit consent model, so a route in one namespace cannot point at a service in another without permission
  • Portability across conformant implementations for the routing layer, which is genuinely more than Ingress ever offered
  • An extension model where implementation-specific policy attaches to a route through a typed reference rather than a string annotation

What it does not do

  • It does not give you authentication, rate limiting, quotas or transformation; those remain implementation-specific extensions
  • Conformance levels vary, so two implementations that both claim Gateway API support will differ on the parts you care about
  • Policy attachment semantics are the least settled part of the specification, so the vendor-specific portion of your config is still vendor-specific
  • It is more verbose than Ingress for simple cases, and for a cluster with three routes that verbosity is pure cost
  • Adopting it does not remove your existing Ingress resources; you will run both during a migration that takes longer than planned

The practical judgment: if you are choosing today, prefer an implementation with strong Gateway API support even if you keep using Ingress for a while. If you already run something that works and you have no multi-team ownership problem, Gateway API is not urgent.

Where the service mesh already does this

This is the duplication worth catching before you buy anything. If you run a mesh, you probably already have a gateway.

A mesh gives every pod a sidecar or a node-level proxy, establishes workload identity with mutual TLS, and controls traffic between services. That is east-west. It does not, on its own, handle traffic arriving from outside the cluster, which is north-south, so every mesh ships an ingress gateway component: a proxy at the edge, configured by the mesh control plane, that takes external traffic and puts it into the mesh.

So the honest question is not “mesh or gateway” but “one proxy at the edge or two”. The common configurations:

Mesh gateway only. External traffic hits the mesh’s gateway and is routed by mesh configuration. One control plane, one config language, one place to look during an incident. You inherit the mesh’s policy vocabulary, which is strong on routing, retries and mTLS and weak on consumer-facing concerns like API keys and quotas.

Ingress controller only, no mesh. External traffic hits the controller, which routes to services directly. Simplest thing that works, and the right answer for most clusters. You get no workload identity and no encryption between services unless you arrange it separately.

Both, chained. An API gateway at the edge handling consumer identity, keys and quotas, forwarding into a mesh that handles service-to-service identity and reliability. This is a legitimate architecture and the right one when your external API is a product. It is also two proxies in the path and two control planes to operate, so adopt it deliberately.

Both, accidentally. An ingress controller from week one, a mesh installed in month eight, and external traffic now traversing the ingress controller and then a sidecar with overlapping retry and timeout policy in both. This is extremely common. The symptom is a retry storm where one layer’s retries multiply the other’s, or a timeout that fires in the wrong place and produces a confusing error. If you are here, the fix is usually to pick one layer to own timeouts and retries and set the other to pass through.

Needs first-hand data: Deploy the same service behind a mesh ingress gateway, a standalone ingress controller, and both chained, with identical retry and timeout policy. Record added p99 latency per configuration and, more usefully, the observed request amplification when the upstream returns errors. The retry-multiplication effect of chained proxies is widely warned about and almost never quantified.

Per-route config ergonomics: the thing you live with

Feature lists do not capture the difference between these products. What you actually live with is how a per-route policy is expressed and who can change it.

Annotations on an Ingress. Strings in a map, controller-specific, no schema, no validation beyond what the controller does at reconcile time, and a typo produces either silence or a reload failure buried in controller logs. Worse, there is no way to grant a team permission to set a timeout but not to change TLS, because it is all one object. This is the baseline everyone starts from and the reason the rest of this exists.

Custom resources per concern. The gateway ships its own CRDs: one for routing, one for authentication policy, one for rate limits. Typed, validated by the API server, and individually subject to RBAC. This is much better and it is what most of the products below do. The cost is that your configuration is now firmly in that vendor’s shape.

Gateway API resources plus vendor policy CRDs. Routing is portable, policy is not, and the boundary is at least explicit. This is the best available position today and where the category is going.

A config file mounted into the proxy. Some teams run a gateway in Kubernetes without a controller, delivering config as a ConfigMap or a mounted file from a repository. Loses dynamic reconciliation, gains complete clarity about what is running. Underrated for small clusters.

One more ergonomic point that decides real migrations: how the controller handles a bad config. Some controllers reject the invalid resource and keep serving. Some accept it, fail to generate valid proxy config, and stop applying all subsequent updates, so a typo in one team’s route silently freezes routing for everyone. Ask this question in every evaluation, because it is the difference between a bad afternoon and an outage.

Kong Ingress Controller

Kong Ingress Controller homepage

Kong’s Kubernetes deployment runs the Kong data plane with a controller that translates Ingress, Gateway API and Kong’s own CRDs into gateway configuration. The draw is that you get Kong’s plugin ecosystem, meaning authentication, rate limiting and transformation as configuration, applied through Kubernetes resources rather than an admin API.

Pros

  • Full Kong plugin ecosystem available as Kubernetes custom resources, so policy is declarative and reviewable
  • Supports Ingress, Gateway API and native CRDs simultaneously, which makes incremental migration practical
  • Can run entirely without a database in Kubernetes, with the controller as the only source of truth
  • Consumer and credential objects are Kubernetes resources, so API key management fits into existing GitOps flows

Cons

  • Two configuration vocabularies coexist, Gateway API and Kong CRDs, and knowing which to use for a given task takes real learning
  • The features platform teams most want, including RBAC and audit over gateway configuration, remain in the commercial build
  • Nginx and Lua data plane means reload behaviour and memory characteristics differ from the Envoy-based options in ways that surprise people migrating

Best for: Teams that want a full policy-capable gateway in Kubernetes with declarative config, and value plugin breadth over data plane uniformity.

Pricing: Open source controller and gateway have no licence cost. The commercial offering meters on the control plane and the number of data plane nodes or services under management.

Traefik

Traefik homepage

Traefik in Kubernetes is the same binary that serves as a Docker reverse proxy, with a Kubernetes provider that discovers routes from Ingress, its own IngressRoute CRD, or Gateway API. It is the lowest-friction option in this list: one deployment, no external dependency, automatic certificate management, and routing that follows your services without a separate config step.

Pros

  • Lowest operational overhead of any option here, with a single binary and no control plane to run separately
  • Automatic ACME certificate management works with minimal configuration, which removes a recurring chore
  • IngressRoute CRD is typed and considerably more pleasant than annotation soup, and Gateway API support exists alongside it
  • Dashboard gives an immediate, honest view of which routes and middlewares are actually loaded

Cons

  • Policy depth is the shallowest here; there is no consumer model, and quota or key management needs another layer
  • Distributed rate limiting and several authentication middlewares are in the commercial build
  • Middleware plugin system is less mature than the plugin ecosystems of Kong or APISIX, so complex policy chains are awkward

Best for: Clusters that need reliable routing, TLS and light middleware without a platform team to operate anything more involved.

Pricing: Open source proxy has no licence cost. The commercial tier is subscription-based by instance, adding advanced middleware, management interface and support.

Apache APISIX

Apache APISIX homepage

APISIX in Kubernetes runs its data plane with an ingress controller in front of etcd, or in a controller-managed mode where the Kubernetes API is the source of truth. It brings the most complete open source policy feature set of the options here, which is the reason to accept its extra moving parts.

Pros

  • The broadest set of policy features available without a commercial licence, including advanced auth and traffic control plugins
  • Plugin runners allow policy in Java, Go, Python or WASM rather than only Lua
  • Strong protocol coverage including gRPC transcoding and non-HTTP proxying, which matters for mixed workloads
  • Configuration changes propagate through watches, so updates apply quickly without proxy restarts

Cons

  • etcd is production infrastructure you now operate, and its failure modes are not intuitive for teams that have not run it
  • Two deployment topologies with different tradeoffs, and the documentation does not make the choice as clear as it should be
  • Documentation and community answers lag the commercial alternatives, so unusual problems take longer to resolve

Best for: Platform teams that want maximum policy capability without licence cost and can absorb the etcd operational burden.

Pricing: Apache-licensed with no licence cost. Commercial support and managed control planes are available from ecosystem vendors, priced by instance and support tier.

Emissary

Emissary is an Envoy-based ingress controller configured entirely through Kubernetes custom resources, and it was one of the first to take the position that ingress config belongs in typed CRDs rather than annotations. Its Mapping resource remains one of the cleaner expressions of a route in this list.

Pros

  • Mapping CRD is a genuinely clean per-route abstraction, and self-service routing by application teams was a design goal rather than an afterthought
  • Envoy data plane brings mature HTTP/2, gRPC and observability behaviour
  • Rate limiting and authentication are delegated to external services over a defined protocol, so you can implement your own policy backend
  • Long-established in Kubernetes specifically, so the Kubernetes-shaped edge cases are well worn

Cons

  • The external auth and rate limit service model means the useful default implementations are in the commercial product, and running your own is real work
  • Envoy configuration generation is opaque when something goes wrong, and debugging requires reading Envoy config dumps
  • Smaller momentum than the Envoy-based projects with foundation backing, which matters for a component this central

Best for: Teams that want Envoy with a clean self-service routing CRD and are willing to supply their own auth and rate limit services.

Pricing: Open source build has no licence cost. The commercial edge stack is subscription-based, adding managed auth, rate limiting, developer portal and support.

Contour

Contour is a deliberately narrow Envoy ingress controller: it does routing, TLS and traffic splitting well, and it does not try to be an API management platform. Its HTTPProxy CRD supports delegation, where a root proxy owned by the platform team delegates a path prefix to a namespace, which was solving the Gateway API ownership problem before Gateway API existed.

Pros

  • Small, comprehensible scope, which makes it one of the easiest Envoy controllers to actually operate
  • HTTPProxy delegation gives real multi-team route ownership with a clear permission boundary
  • Rejects invalid configuration and reports status on the resource rather than silently breaking the data plane
  • Strong Gateway API support, so it is a reasonable target for teams standardising on the specification

Cons

  • No policy layer at all: no authentication, no rate limiting, no consumer model, by design
  • You will need another component for anything beyond routing, so it is a piece of an architecture rather than a complete answer
  • Smaller ecosystem of examples and integrations than the larger projects

Best for: Clusters that want a dependable Envoy routing layer with clean multi-tenancy, where policy is handled elsewhere in the stack.

Pricing: Open source with no licence cost; support comes from the Kubernetes distribution or vendor you already run.

Istio Gateway

Istio’s gateway is the mesh’s north-south entry point, configured by the same control plane as everything else in the mesh. Choosing it is not really a gateway decision; it is a decision to let the mesh own the edge as well as service-to-service traffic. If you already run Istio, this is usually the right answer, and if you do not, installing Istio to get a gateway is a large commitment for a small benefit.

Pros

  • One control plane and one configuration vocabulary for both external and internal traffic, so there is one place to look during an incident
  • Traffic management is the strongest here: weighted splits, mirroring, fault injection and retry policy are all first-class
  • Workload identity and mutual TLS extend from the edge all the way through, which is a genuine security property rather than a feature bullet
  • Gateway API support is well developed and is now the recommended configuration path

Cons

  • Installing Istio purely for an edge gateway means operating a mesh control plane, which is a substantial ongoing commitment
  • Consumer-facing concerns like API keys, quotas and developer onboarding are not its model, so an external API product needs another layer
  • The interaction between mesh sidecar policy and gateway policy produces failure modes that are hard to reason about, especially around retries and timeouts

Best for: Organisations already running Istio, who want the edge governed by the same policy engine as the mesh.

Pricing: Open source with no licence cost. Commercial distributions and managed control planes are available from several vendors, priced by cluster or workload count.

Envoy Gateway

Envoy Gateway is the Envoy project’s own control plane for Kubernetes, built around Gateway API as the primary configuration model rather than an afterthought. Its value proposition is that it is the reference way to run Envoy at the edge without adopting a mesh or a vendor control plane, with policy attached through typed extension resources.

Pros

  • Gateway API is the native configuration model, not a translation layer over something else
  • Backed by the Envoy project itself, which reduces the risk that the control plane and data plane drift apart
  • Extension policy resources cover authentication, rate limiting and traffic control with typed CRDs
  • Considerably simpler to operate than a full mesh while using the same proven data plane

Cons

  • Younger than the alternatives, so fewer people have operated it through a genuine incident and community answers are thinner
  • Policy feature set is narrower than the mature gateways, and the gaps move quickly enough that documentation lags
  • Rate limiting depends on a separate rate limit service and a Redis backing it, which is more infrastructure than the single-binary options

Best for: Teams standardising on Gateway API who want Envoy at the edge without adopting a mesh or a commercial control plane.

Pricing: Open source with no licence cost; the operational cost is the control plane, the rate limit service and its datastore.

NGINX Ingress

NGINX Ingress is the incumbent, and it is on most clusters because it was the default when the cluster was built. It routes reliably, it is extremely well understood, and its configuration model is annotations that expand into an Nginx configuration file. That model is both why it is everywhere and why teams eventually leave it.

Pros

  • The most widely deployed option, so nearly every problem you hit has been hit publicly before
  • Nginx behaviour under load is well understood and predictable, with decades of operational knowledge available
  • Minimal moving parts: a controller and an Nginx process, with no external control plane or datastore
  • Snippet annotations allow raw Nginx configuration for cases nothing else covers

Cons

  • Annotation-based configuration has no type safety and no per-field RBAC, and raw config snippets are a genuine multi-tenant security concern that most clusters end up disabling
  • Configuration reloads regenerate the whole Nginx config, so one team’s bad annotation can break routing for everyone in the cluster
  • No consumer model, so per-customer quotas and key management require another layer entirely

Best for: Single-team clusters that need dependable routing, already understand Nginx, and have no multi-tenancy or per-consumer policy requirement.

Pricing: Open source with no licence cost. A commercial variant built on the commercial Nginx product is sold separately by subscription.

How to choose

Start by writing down what already terminates external traffic. If it is a mesh gateway, your decision is whether to add a layer, not which layer to install. If it is an ingress controller with fifty annotations, your decision is whether to migrate the annotations to typed resources or replace the controller.

Then decide whether you need a consumer model. Per-customer API keys, quotas and usage reporting are the dividing line in this list. Contour, NGINX Ingress and the free Traefik tier do not give you one. Kong, APISIX, Emissary and the commercial tiers do. If you need it and pick something that lacks it, you will build it badly in application code.

Then decide who owns routing configuration. One team owning everything means annotations are survivable and Gateway API is optional. Multiple teams self-serving routes means you need a real ownership boundary, which is Gateway API, Contour’s delegation, or a controller with per-namespace CRDs and cluster RBAC.

Then choose on control plane operational complexity, not data plane performance. Five of the eight options here are Envoy. The data planes are close enough that the difference will not decide anything. What differs is how many processes you operate, what external state exists, and how the system behaves when a config is invalid.

OptionData planeConfig modelConsumer modelExtra infrastructure
Gateway API (standard, not a product)n/aTyped routing CRDs with role separationNon/a
Kong Ingress ControllerNginx and LuaGateway API plus Kong CRDsYesNone required
TraefikTraefikIngress, IngressRoute CRD, Gateway APINo in open sourceNone
Apache APISIXNginx and LuaAPISIX CRDsYesetcd
EmissaryEnvoyMapping CRDsExternal servicesAuth and rate limit services
ContourEnvoyHTTPProxy, Gateway APINoNone
Istio GatewayEnvoyIstio CRDs, Gateway APINoMesh control plane
Envoy GatewayEnvoyGateway API plus policy CRDsPartialRate limit service and Redis
NGINX IngressNginxIngress annotationsNoNone

Whatever you choose, treat the gateway’s telemetry as part of your cluster observability rather than a separate silo, which is a recurring theme in both Kubernetes APM and Kubernetes log management. And keep the gateway configuration in the same repository and reconciliation flow as the rest of your cluster, which is the argument made at length in GitOps tooling.

Needs first-hand data: Apply a deliberately invalid route resource to each candidate controller in a cluster serving live traffic, and record three things: whether existing routing continues to work, whether subsequent valid updates still apply, and how clearly the failure is reported on the resource status. This single test separates these products more sharply than any feature comparison and nobody publishes the results.

Frequently asked questions

Should I migrate from Ingress to Gateway API?

If several teams edit routing configuration, yes, because the role separation is the point and it solves a problem annotations cannot. If one team owns everything and your current setup is stable, it is not urgent. Either way, choose new components with strong Gateway API support so the migration stays available to you.

Do I need both an ingress controller and a service mesh?

Only if you need what each uniquely provides. The mesh gives workload identity and service-to-service encryption. A gateway gives consumer identity, keys and quotas. If you need only the first, use the mesh gateway for the edge too. If you need only the second, skip the mesh. Running both is defensible and means operating two proxies with overlapping retry and timeout policy, which you must configure deliberately.

Is Envoy always the better data plane?

Not always, but it is the safer default. Envoy has strong HTTP/2 and gRPC support, a well-specified configuration model, and excellent built-in observability. Nginx-based gateways are lighter in memory and extremely well understood operationally. The real difference is reload behaviour: Envoy updates configuration incrementally over xDS while Nginx-based controllers regenerate and reload, which is why a single bad route can have a wider blast radius on the Nginx side.

Can I use a cloud provider load balancer instead?

For TLS termination and path routing, often yes, and it is one fewer component to run. What you give up is anything per-consumer, most transformation capability, and the ability to express routing as Kubernetes resources reviewed alongside your deployments. Many clusters end up with both: a cloud load balancer in front for the public address and TLS, and a controller inside for routing.

How do I rate limit per customer in Kubernetes?

Not with node-local counters, because with more than one replica each replica enforces its own limit and the effective limit is the configured value multiplied by replica count. You need a shared counter, which means a rate limit service and a datastore, or a gateway with distributed rate limiting built in. This is the most common thing teams get wrong in this category, and it is covered properly in rate limiting solutions.