Buyer’s Guide

Best API Gateways

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • api
  • gateway
  • infrastructure
  • proxy

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

Most gateway comparisons start with a feature matrix. That is the wrong end of the problem, because every gateway in this category does authentication, rate limiting, request transformation and observability. The rows all get ticked and nothing gets decided.

The decision is actually about what you are putting in the request path. A gateway is a process that every single call to your API now traverses. It adds latency to every request, it adds a component that can be down while all your services are healthy, and it adds a configuration surface where a mistake takes out traffic that was previously fine. In exchange it gives you one place to enforce policy instead of fifteen inconsistent middleware stacks.

That trade is almost always worth making past a certain number of services. But you should make it knowingly, and you should pick the gateway whose failure modes you can actually operate at 3am, not the one with the longest plugin list.

So the mechanism comes first here. What happens to a request inside a gateway, where the latency goes, what state the gateway needs and what happens when that state is unavailable, and how configuration reaches the data plane. Then the products, evaluated against that.

Key takeaways

  • A gateway adds a network hop and a policy chain to every request. The latency is usually small and the new failure domain is not, and the second one is what should drive your choice.
  • The most important architectural question is what state the data plane needs to serve a request. A gateway that must reach a database or Redis on the request path has coupled your API availability to that store.
  • Declarative configuration applied from a repository is a materially different operational model from a live admin API, and it is the single biggest predictor of whether gateway changes cause outages.
  • Plugin and policy code is the part that does not migrate. Everything else in this category is roughly portable; custom Lua, custom policy XML and vendor-specific authorizers are not.

What a gateway does to a request

Strip the marketing and a gateway does six things in order, and each one costs something.

Accept and terminate. The connection arrives, TLS is terminated, and the request is parsed. This is where your HTTP/2 and HTTP/3 support lives, where connection limits apply, and where a slow client stops being your application’s problem and starts being the gateway’s. Termination at the edge is one of the quietly valuable things a gateway does: your upstream services stop having to think about certificate rotation or slowloris behaviour.

Match a route. The gateway maps method, host, path and sometimes headers to a route object, and the route names an upstream. Route matching is usually a radix tree or similar, so the cost is small and roughly independent of route count. What is not independent of route count is the memory to hold them and the time to rebuild the matcher when configuration changes, which matters once you have thousands of routes.

Run the policy chain. This is the part that costs real time. Each plugin or filter runs in sequence: validate the API key, check the rate limit, verify the JWT signature, rewrite the path, inject headers, emit a metric. Some of these are pure CPU, some make network calls. The plugins that make network calls are the ones that determine your added latency, and there are usually two: the rate limiter talking to a shared counter, and the authenticator talking to an identity provider or introspection endpoint.

Select an upstream and forward. Load balancing across upstream instances, connection pooling, and the retry and timeout policy. Connection pooling to upstreams is another genuine win: your services see a small number of long-lived connections from the gateway rather than a churn of client connections.

Handle the response. Transform, cache, compress, add CORS headers. Response transformation is where buffering happens, and buffering is where streaming responses and large payloads get expensive in memory.

Emit telemetry. Access logs, metrics, traces. Almost always configurable, and almost always the first thing to become a bottleneck when someone turns on full request and response body logging in production.

The sum of steps three and six is your added latency. Steps one, two, four and five are usually a small fixed cost. If a gateway’s benchmark looks great and yours looks bad, the difference is almost always a plugin making a synchronous network call that the benchmark did not run.

Needs first-hand data: Run the same upstream behind each gateway with three configurations: bare proxy, proxy plus JWT validation, and proxy plus JWT validation plus distributed rate limiting. Record added p50, p99 and p99.9 at a fixed concurrency. The delta between configuration one and three is the number that actually matters, and it is not published anywhere.

The state question decides more than the feature list

Here is the architectural distinction that separates these products better than anything on a comparison page: what does the data plane need in order to serve a request, and what happens when it cannot reach it.

Stateless data plane, config pushed ahead of time. The gateway holds its entire configuration in memory. It was pushed there by a control plane or loaded from a file at startup. Serving a request requires nothing external. If the control plane dies, the gateway keeps serving the last configuration it received, forever, which is exactly what you want during an incident. KrakenD and Traefik sit here by design, and Kong’s declarative mode and Envoy’s xDS model both aim at it.

Data plane reads config from a store on the request path. The gateway looks up routes, consumers or credentials in a database or cache as requests arrive, usually with a local cache in front. This is more flexible, because a config change is visible immediately with no push, and it is strictly worse for availability. A slow database is now a slow API. Cache warming behaviour after a restart becomes something you have to think about.

Data plane needs shared state for correctness. Even a stateless configuration model needs shared state for two things: distributed rate limiting and revoked-credential lookups. A rate limit counter that is local to each node is not really a rate limit once you run more than one node, so there is usually a Redis or a dedicated rate limit service involved. This is the most commonly underestimated dependency in the whole category, and it deserves its own decision, covered in rate limiting solutions.

Ask every vendor one question: if the store is unreachable, does the gateway fail open or fail closed, and is that configurable per policy? The right answer for rate limiting is usually fail open, because rejecting all traffic because you cannot count it is worse than briefly not counting. The right answer for authentication is fail closed. A gateway that applies the same policy to both is one you will have to work around.

Configuration delivery is the operational model

The second thing that separates these products is how configuration reaches the data plane, and this predicts outages better than any performance number.

A file in a repository, applied by CI. The config is reviewed, diffed, versioned and rolled back like code. Someone can see what changed and when. This is the model that makes gateways boring, and boring is the goal.

An admin API with a database behind it. Powerful and immediate, and it means gateway state can drift from anything you have written down. The failure pattern is that someone fixes an incident at 2am by POSTing to the admin API, the fix is never written back to the repo, and the next full redeploy silently reverts it.

Kubernetes custom resources. Config is Kubernetes objects, reconciled by a controller. This gets you the repository model for free if you already practice GitOps, plus RBAC from the cluster. It also inherits every Kubernetes failure mode and gives per-route config a YAML shape that gets verbose fast. This is a separate shortlist, covered in API gateways for Kubernetes.

If you take one thing from this article: prefer a gateway whose primary configuration model is declarative and file-based or resource-based, even if it also offers an admin API. The admin API is convenient and it is how gateway configuration becomes undocumented.

Kong Gateway

Kong Gateway homepage

Kong is the default answer in the self-hostable tier, built on Nginx and OpenResty with plugins written in Lua. It runs in three quite different modes: database-backed with an admin API, declarative from a config file with no database, and as a data plane attached to a hosted control plane. Those modes have genuinely different operational characteristics, so “we run Kong” tells you much less about a team’s setup than it sounds like it should.

Pros

  • Largest plugin ecosystem in the self-hostable tier, so most policy requirements are configuration rather than code
  • Declarative mode removes the database from the request path entirely and turns gateway config into a reviewable artefact
  • Custom plugins in Lua are genuinely approachable, and the plugin development experience is better than most competitors
  • One data plane serves standalone, Kubernetes and hosted control plane deployments, so the deployment choice stays reversible

Cons

  • The features a platform team wants most, meaning RBAC, workspaces and audit logging, sit in the commercial build rather than the open source one
  • Database-backed mode couples API availability to Postgres availability, and teams adopt it without noticing they made that trade
  • Plugin ordering is a real operational skill; two plugins that each work fine can misbehave in combination and the logs will not tell you why

Best for: Teams who want a self-hosted gateway with the broadest plugin coverage and will commit to declarative configuration rather than the admin API.

Pricing: Open source build carries no licence cost. The commercial offering meters on the hosted control plane plus the number of data plane nodes or services under management, with enterprise plugins gated by tier.

Tyk

Tyk homepage

Tyk is a Go gateway with Redis for distributed state and a commercial dashboard on top. The architectural difference from Kong is where the commercial line falls: most of the gateway’s policy capability is in the open source build, and the dashboard, portal and multi-tenancy are what you pay for. That makes the open source experience unusually complete if you are willing to drive it through its API.

Pros

  • Single Go binary is a small operational footprint compared to an Nginx and Lua stack
  • Policy features including rate limiting, quotas and key management are in the open source gateway rather than behind the wall
  • Multi-tenancy model is stronger than most competitors, which matters when you serve several distinct consumer organisations
  • Native support for API key and quota semantics rather than treating them as a plugin afterthought

Cons

  • Redis is a hard dependency for keys, quotas and distributed rate limits, so Redis availability is API availability
  • Without the commercial dashboard, day-to-day management is API-driven and noticeably less pleasant than the demo suggests
  • Smaller plugin ecosystem than Kong or APISIX, so unusual integrations mean writing Go middleware or a gRPC plugin service

Best for: Teams who want a lightweight self-hosted gateway with first-class API key and quota handling, and already run Redis they trust.

Pricing: Open source gateway has no licence cost. Commercial tiers meter on dashboard, portal and the number of gateway nodes or managed APIs, with self-managed and hosted options.

KrakenD

KrakenD homepage

KrakenD is the outlier in this list and the most interesting one architecturally. It is a stateless gateway that reads a single immutable configuration file at startup and holds everything in memory. There is no database, no admin API and no runtime configuration change: to change config you deploy a new instance. It is also built around API composition, aggregating several backend calls into one response, which is a different product thesis from policy enforcement.

Pros

  • Genuinely stateless with no runtime dependencies, so the gateway cannot be taken down by a store it cannot reach
  • Immutable config means the running configuration is always exactly what is in your repository, with no drift possible
  • Backend aggregation and response filtering are first-class, which removes a class of backend-for-frontend services entirely
  • Small resource footprint and fast startup, which makes it a good fit for autoscaling and ephemeral environments

Cons

  • No runtime configuration changes at all, so every route addition is a deployment, which some teams experience as discipline and others as friction
  • Distributed rate limiting needs external help, because the stateless design has no shared counter by default
  • Smaller ecosystem and a configuration format that gets large and repetitive for estates with many endpoints

Best for: Teams that want a gateway they can treat as an immutable artefact, particularly where aggregating several backends into one client-facing response is a real requirement.

Pricing: Open source build has no licence cost. The enterprise build is subscription-based and adds the management interface, additional policy modules and support.

Apache APISIX

Apache APISIX homepage

APISIX is the other significant Nginx and Lua gateway, differing from Kong mainly in its control plane: configuration lives in etcd and is watched by the data plane, so changes propagate without a restart and without a request-path database read. It is a genuine Apache Software Foundation project rather than an open-core product with a foundation-shaped wrapper, and the feature split reflects that.

Pros

  • Configuration changes propagate through etcd watches, so updates are near-immediate without a request-path store lookup
  • Very complete open source feature set, including features that are commercial in comparable products
  • Plugin runners let you write plugins in Java, Go, Python or WASM rather than only Lua
  • Strong protocol coverage including gRPC transcoding, MQTT and TCP/UDP proxying

Cons

  • etcd is now production infrastructure you own, and etcd operational problems are not intuitive if you have not run it before
  • Documentation and ecosystem maturity lag the commercial competitors, and answers are more often in issue threads than in docs
  • The commercial offering around it is smaller, so the support story is weaker if procurement requires a vendor with a large support organisation

Best for: Teams who want the most complete open source feature set and are comfortable operating etcd as part of their platform.

Pricing: Apache-licensed with no licence cost. Commercial support and a managed control plane are available from vendors in the ecosystem, priced by instance and support tier.

Traefik

Traefik homepage

Traefik started as a reverse proxy that discovers its own configuration from the platform it runs on, and that remains its distinguishing property. Point it at Docker, Kubernetes or a service registry and routes appear as services appear, with no separate configuration step. It has grown genuine gateway features since, but the auto-discovery model is still why teams pick it.

Pros

  • Provider-based auto-discovery means routes track your actual deployments rather than a parallel config file that drifts
  • Automatic certificate management including ACME is built in and works with very little configuration
  • Single Go binary, no external dependencies for basic operation, and a low operational learning curve
  • The same binary serves as a Docker reverse proxy, a Kubernetes ingress controller and a Gateway API implementation

Cons

  • Policy depth is shallower than the dedicated gateways; advanced auth, quota and transformation needs push you toward plugins or another layer
  • The middleware plugin system is less mature than Kong or APISIX, and complex policy chains are harder to express
  • Advanced features including distributed rate limiting and some authentication middleware sit in the commercial build

Best for: Teams who want routing, TLS and basic policy to follow their deployments automatically, and whose policy requirements are moderate.

Pricing: Open source proxy has no licence cost. The commercial offering is subscription-based, adding advanced middleware, a management interface and support, scaled by instance count.

Zuplo

Zuplo homepage

Zuplo is the newest architectural bet here: a gateway that runs on edge runtimes, where policies are TypeScript modules rather than plugin configuration, and where the whole configuration is a Git repository deployed like an application. The consequence is that the gateway is deployed close to callers by default and that custom policy is written in a language your team already uses.

Pros

  • Policies are TypeScript, so custom logic is written and tested with the tooling your application developers already have
  • Configuration is a Git repository with preview deployments per branch, which brings normal software review to gateway changes
  • Edge deployment puts the gateway near callers without you operating points of presence
  • Developer portal and API key management are integrated rather than a separate product tier

Cons

  • Edge runtimes are not full Node, so the libraries available to your policy code are constrained in ways that surface late
  • Managed only, so there is no self-hosted escape hatch if data residency or air-gapped operation becomes a requirement
  • Younger product with a smaller ecosystem, and fewer people who have operated it through a bad day

Best for: Product teams shipping a public API who want gateway policy to live in their own repository in TypeScript, and who have no self-hosting requirement.

Pricing: Usage-based on requests with tiered plans, and separate metering for additional environments and portal features.

Gravitee

Gravitee homepage

Gravitee ships a gateway, a management API, a developer portal and an access management component, most of it open source. Its distinguishing technical bet is event-native APIs: Kafka, MQTT and server-sent events are first-class backends, and it can expose an event stream as a REST or WebSocket API rather than requiring you to build a bridge service.

Pros

  • Event and streaming protocol support is architectural rather than a bolt-on, which is rare in this category
  • The open source build includes the portal and management console, not just the proxy
  • Self-hostable end to end, so data residency and air-gapped deployments are achievable
  • Policy studio makes complex policy chains visible and editable without writing plugin code

Cons

  • More components to operate than a single-binary gateway, since a full deployment is several services plus a datastore
  • Smaller community than Kong or APISIX, so unusual problems are more often yours alone to solve
  • The line between open source and enterprise features moves, so verify your tier rather than trusting a feature page

Best for: Teams that need a self-hostable gateway and portal together, particularly where Kafka or MQTT streams are part of the API surface.

Pricing: Open source build has no licence cost. The enterprise build is subscription-based, scaled by environments and gateway instances under management.

AWS API Gateway

AWS API Gateway removes the operational question entirely: there are no nodes, no scaling decisions and no capacity planning. What you get instead is a policy model that is thinner than the dedicated gateways and a pricing model that scales linearly with request count. It exists in distinct flavours with materially different capabilities, and choosing the wrong one is the most common mistake teams make with it.

Pros

  • No infrastructure to run, which for a small team is worth more than any feature comparison
  • IAM, Cognito and resource policy integration makes authentication configuration rather than deployment
  • Direct integrations let you front Lambda, Step Functions, SQS or DynamoDB without writing a proxy service at all
  • Regional and edge deployment, throttling and caching are all managed concerns

Cons

  • Per-request pricing makes high-volume, low-value traffic structurally expensive compared to a node you run yourself
  • Policy model is thin; anything beyond the built-ins becomes a Lambda authorizer, which adds its own latency and cold start behaviour
  • Deepest configuration lock-in of any option here, with hard payload size and timeout ceilings that surface late and cannot be tuned

Best for: Teams already committed to AWS with moderate request volume and mostly serverless backends, who want zero gateway operations.

Pricing: Per-request metering with different rates by gateway type, plus data transfer, plus optional caching billed by cache size per hour.

Apigee

Apigee homepage

Apigee is an API program suite that includes a gateway, and reading it as a gateway comparison misses the product. It assumes an API with external consumers, a lifecycle with approvals, a developer portal, tiered products and someone accountable for adoption. The proxy is competent; the surrounding program machinery is what you are buying.

Pros

  • The most complete developer portal, API product and monetisation model in the category
  • Proxy revisions are versioned, promotable artefacts rather than live configuration edits, which suits regulated change control
  • Analytics by developer, app and API product is the reporting an external API program actually needs
  • Policy model covers heavy enterprise requirements including message-level security and complex mediation

Cons

  • Large conceptual surface: proxies, products, developers, apps, environments and revisions are separate objects you must model correctly
  • Policy is expressed in a vendor-specific form, which is the least portable configuration of any option here
  • Strongly Google Cloud oriented, with an enterprise sales motion that makes casual evaluation slow and poor value for purely internal APIs

Best for: Organisations where the API is a commercial product with external developers, tiered access and formal change control.

Pricing: Subscription tiers gated by API call volume, with separate entitlements for environments, monetisation and advanced security, sold on annual commitment.

Azure API Management

Azure API Management is the Microsoft equivalent of the program suite, bundling gateway, developer portal and policy management into one service, with a self-hosted gateway option that lets you run the data plane outside Azure while keeping the control plane managed. That hybrid mode is its most distinctive feature and the reason it appears in shortlists where Azure is not the primary cloud.

Pros

  • Self-hosted gateway mode runs the data plane in your own environment against a managed control plane, which suits hybrid and data residency constraints
  • Policy expressions are powerful and cover complex mediation, transformation and conditional logic
  • Built-in developer portal and product and subscription model without buying a separate tier
  • Deep Entra ID integration makes enterprise authentication straightforward if you are already a Microsoft shop

Cons

  • Policy is vendor-specific XML, which is expressive and completely non-portable
  • Tier boundaries are sharp, and features like virtual network integration and multi-region only appear at higher tiers
  • Provisioning and scaling operations on some tiers are slow enough to be a planning consideration rather than an afterthought

Best for: Organisations already standardised on Azure and Entra ID, especially those needing a managed control plane with data planes running in their own environments.

Pricing: Tiered by instance with capacity units, with per-tier feature gating and consumption-based options for lower volumes.

How to choose

Work through these in order. Each one eliminates candidates, and the order matters because the early ones are constraints while the later ones are preferences.

Constraint one: can it be managed? If data residency, air-gapped operation or an on-premise requirement exists, the managed-only options are gone and you are choosing among self-hostable data planes. If nobody cares where it runs and your team is small, managed removes real work.

Constraint two: what is the primary deployment target? Kubernetes changes the shortlist substantially, because the Gateway API, ingress controllers and your service mesh already overlap with this job. VMs favour single-binary gateways. Serverless favours the cloud provider’s own.

Constraint three: what does the request path depend on? Write down, for each candidate, every external system a request touches. Then decide whether you are willing to run each of those at your API’s availability target. This eliminates more candidates honestly than any feature comparison.

Preference one: how is configuration delivered? Prefer declarative and reviewable. Treat a live admin API as an emergency tool, not the primary workflow.

Preference two: how much custom policy will you write? If the answer is none, the plugin ecosystem size is what matters. If the answer is a lot, the language and testing story for custom policy matters much more, and it is the part of your investment that will not migrate.

Preference three: the pricing unit. Per request punishes chatty internal traffic. Per node punishes running many small isolated gateways. Per API product punishes fine-grained service decomposition. Project twelve months ahead, not today.

Needs first-hand data: For one real service, measure the request-path dependency cost directly. Configure distributed rate limiting against a shared store, then induce latency on that store and record what happens to gateway p99 and error rate under each candidate’s fail-open and fail-closed settings. The failure behaviour differs sharply between products and is documented vaguely everywhere.

Two further notes before procurement. First, whatever you pick, put the gateway’s own metrics into the same place as your service metrics, because a gateway that is observed separately from the services behind it produces incidents where two teams look at two dashboards and disagree. External API monitoring matters too, since the gateway cannot report its own unavailability. Second, decide early where token validation happens; the API authentication choice constrains the gateway choice more than the reverse.

Frequently asked questions

How much latency does an API gateway actually add?

For a bare proxy with no policy, the added latency is small and dominated by one extra network hop plus TLS termination. The number changes entirely once policy runs. A JWT signature verification is CPU work measured in fractions of a millisecond. A token introspection call to an identity provider, or a rate limit check against a remote Redis, is a network round trip on every request, and that is what you will actually feel. Measure with your policy chain enabled, never with a bare proxy.

Should the gateway do authentication or should each service?

The gateway should reject anything with an invalid or missing credential, so unauthenticated traffic never reaches a service. Services should still make authorisation decisions, because only they know what a caller may do with a specific resource. The anti-pattern is pushing fine-grained authorisation into gateway configuration, which puts business rules in proxy config where nobody tests them.

Can I run an API gateway without a database?

Yes, and for most teams you should. Kong in declarative mode, KrakenD, Traefik and Envoy-based gateways all serve requests from in-memory configuration. You will still typically need shared state for distributed rate limiting, but that is a narrower dependency with an acceptable fail-open behaviour, unlike a database that routing itself depends on.

Is an API gateway the same as a load balancer?

No, though the boundary keeps moving. A load balancer distributes connections and increasingly does path routing and TLS termination. A gateway adds a consumer identity model: it knows which API key or token made the call, applies per-consumer quotas, and reports traffic by consumer. If you never need to answer “which customer sent these requests”, a load balancer may genuinely be enough.

Do I need a gateway if I already run a service mesh?

Usually yes, for the north-south edge. A mesh secures and observes traffic between workloads you control, using workload identity. A gateway handles traffic from callers you do not control, where the identity is a customer, the credential is an API key or token, and quotas and monetisation apply. Several mesh projects ship a gateway component precisely because the mesh alone does not cover this.

What is the hardest part of migrating between gateways?

Custom policy code, in every case. Routes, upstreams and TLS configuration translate mechanically. Custom Lua plugins, Apigee policy XML, Azure policy expressions and Lambda authorizers do not translate at all, and they are usually where the business-specific behaviour accumulated. Bound how much of that you write, and keep the intent documented separately from the implementation.