Every comparison of these three lists the same capabilities down the left-hand side — multi-provider routing, fallbacks, caching, rate limits, spend tracking — puts ticks in most of the boxes, and concludes that they are broadly similar with different pricing. That table is accurate and completely useless, because all three genuinely do those things, and none of them will be the reason you regret your choice.
The reason you will regret your choice is structural. These are three different bets about where an LLM gateway belongs in an organisation. LiteLLM bets it is infrastructure your team runs. Portkey bets it is a managed platform you buy, and since becoming part of Palo Alto Networks that platform’s centre of gravity is security and governance. Kong bets it is not a new thing at all — that LLM traffic is another protocol on the API gateway you already operate.
Pick the wrong bet and the symptoms show up eighteen months later: a platform team maintaining a Python proxy nobody wants to own, a procurement conversation about a security vendor’s platform you did not intend to buy into, or a second control plane duplicating governance you already had.
So the useful question is not which has more features. It is which of those three bets matches how your organisation actually works.
Key takeaways
- All three do routing, fallback, key custody, caching and spend visibility. Stop comparing on those; they are table stakes and they will not differentiate your outcome.
- The three real axes are who operates it, whether LLM traffic is a special case or just traffic, and what happens to governance as agents and MCP servers multiply.
- Portkey is now Prisma AIRS AI Gateway inside Palo Alto Networks. For some buyers that is the reason to choose it; for others it changes the decision from a tooling choice to a security-platform choice.
- Kong’s advantage is organisational before it is technical: one gateway of record, one policy layer, one team that already knows it.
What all three do identically
Get these off the table first, because arguing about them wastes the evaluation.
Multi-provider routing behind one endpoint. All three accept an OpenAI-shaped request, resolve it to a provider and model from your configuration, and return a normalised response. Your application code calls one base URL with one key regardless of whether the request ends up at a hosted frontier model or your own inference server.
Fallback and load balancing. All three let you define an ordered list of targets so a failure or rate-limit response on the primary flows to a secondary, and all three can spread traffic across multiple deployments of the same model — several regions of one provider, or several providers serving equivalent models. The configuration syntax differs, the concept does not.
Credential custody and virtual keys. All three hold the real provider credentials and issue their own keys to applications, teams and users. That is the single most valuable thing a gateway does, and the one that pays for itself the first time you rotate a leaked key without touching a deployment. Provider credentials stop being scattered across environment variables in a dozen services.
Caching, and token-aware rate limits. All three cache responses on a request hash, offer semantic caching in some form, and understand that requests-per-minute is the wrong unit for LLM traffic. Cache hit-rate economics are a property of your traffic, not of the gateway.
Spend visibility and budgets. All three attribute token usage and estimated cost to a key, a team or a tag, and all three can refuse a request that would breach a budget. The numbers are always an approximation assembled from provider token counts and a price table the gateway maintains — in every product. Treat differences here as reporting polish, not accuracy.
If your shortlist is these three, every one of those requirements is satisfied. What follows is what is not.
Axis one: who operates it
This is the axis that decides most of these evaluations, and it is a staffing question dressed up as a technology question.
LiteLLM puts the operational load on you. You run the proxy tier, a Postgres holding keys, budgets, spend and optionally request payloads, and a Redis holding rate-limit counters and the cache. You own its uptime — and because every LLM feature routes through it, its uptime is your product’s uptime. You own upgrades on a fast-moving project, and the retention policy for a database that fills with prompts. In exchange you own the data, the deployment location and the exit.
Portkey removes that load and adds a vendor relationship. The managed gateway is somebody else’s on-call rotation, scaling problem and Postgres. What you take on instead is a dependency on their availability, a data path that includes their infrastructure unless you deploy otherwise, and a procurement relationship. For most teams shipping product features that trade is correct: a two-person team should not be operating a three-tier stateful service in the critical path of their main feature.
Kong moves the load to a team that already carries it. This is the underrated option. If a platform team already runs Kong for your REST APIs, adding AI plugins creates no new operational owner, alerting surface, upgrade cadence or on-call rotation — the marginal cost is a plugin configuration. If you do not already run Kong, this advantage evaporates entirely.
The failure mode on this axis is one I have seen in adjacent categories often enough to predict: a product team self-hosts a proxy because it was quick, it becomes load-bearing, the person who set it up changes teams, and eighteen months later nobody will upgrade it because nobody understands the config. Self-hosting is a fine decision with a named owner and a bad one without. The self-hosted gateway operations piece covers what that ownership involves.
Needs first-hand data: Cost your own operational load honestly. Track hours spent on the gateway — deploys, upgrades, incidents, config changes, capacity work — for one quarter after a self-hosted deployment stabilises. Compare that number, at your loaded engineering cost, against a managed gateway quote at your request volume. Most teams have never done this arithmetic and are surprised in both directions.
Axis two: is LLM traffic special, or is it just traffic
This axis sounds abstract and has very concrete consequences.
LiteLLM and Portkey both treat LLM traffic as its own domain. Their abstractions are model-shaped: model lists, provider mappings, token counting, prompt templates, semantic caching, guardrail hooks on completions. That focus is why their provider coverage and LLM-specific features are ahead.
The cost is two control planes. Your REST APIs are governed by whatever you already had — authentication, rate limits, WAF rules, audit logs, network policy, a service catalogue. Your LLM calls are governed by the gateway, with its own policy model, key namespace, logs and idea of what a tenant is. Two places to check when a security review asks who can call what. Two places to update when a policy changes. Two things to monitor.
Kong treats LLM traffic as another protocol on the gateway of record. The same route and plugin model that governs /api/v2/orders governs /v1/chat/completions. Authentication, authorisation, rate limits, logging destinations and network policy are the ones you already wrote; the AI plugins add provider proxying, model load balancing, prompt guarding, semantic caching and request transformation on top of a substrate that already knew how to be a gateway.
The cost is depth. A gateway whose entire product is LLM traffic ships model-specific features sooner and further: newer providers, richer prompt management, finer cost attribution.
Resolve this axis by asking who answers the audit question. If someone in your organisation must be able to say “here is every external call our systems make and who is allowed to make it”, a second ungoverned control plane is a problem you will pay for. If nobody asks that yet, the specialist’s depth is worth more today.
Axis three: what happens as agents and MCP servers multiply
This is the axis most evaluations skip because it is about a problem the team does not have yet. It is also the one most likely to make your choice look prescient or naive in two years.
Right now, in most organisations, LLM traffic is a small number of call sites: a chat feature, a summarisation job, an embedding pipeline. Governing that is easy in any of the three.
What changes the shape is agents. An agent does not make one model call — it makes a variable number, decides at runtime which tools to invoke, and those tools are increasingly MCP servers reaching internal APIs, databases and third-party services. Then agents start calling other agents. The graph stops being “our app calls a model” and becomes “an autonomous process calls a model, which decides to call a tool, which calls a service, which invokes another agent.”
The governance questions that follow are not model questions:
- Which agent is allowed to invoke which MCP server, and who approved that?
- What is the audit record when an agent takes an action with a side effect, and does it link back to the human or system that initiated the chain?
- How do you rate-limit or budget a process whose call count is decided by a model at runtime rather than by your code?
- When a tool call goes wrong, where is the trace that shows the whole chain rather than one hop of it?
Kong’s product framing extends explicitly to this: governing LLM, MCP and agent-to-agent traffic through the same gateway and policy layer. The architectural claim is coherent — if agent traffic ends up looking like a service mesh problem, the thing that governs it should be the component that already governs service traffic.
Portkey’s answer runs through the security-platform side of the same problem: an AI gateway inside a security vendor’s portfolio, where runtime AI security, policy enforcement and threat inspection are the organising idea rather than a plugin on a proxy.
LiteLLM’s answer is that it is a proxy, and governing the wider agent graph is something you assemble around it.
None of those is wrong. They are different bets about which discipline ends up owning agent governance: the platform team, the security team, or the application team.
MCP and A2A: standards, not products
Two of the things this comparison turns on are protocols, and it is worth separating them from products before the product blocks.
MCP (Model Context Protocol) is an open specification for how a model-driven application discovers and calls tools and data sources. An MCP server exposes tools; an MCP client — an agent, an IDE, a chat application — discovers and invokes them over a defined protocol. It is why a tool written once can be used by many agents.
A2A (agent-to-agent) is the equivalent idea one level up: a protocol shape for agents to advertise capabilities and delegate work to each other, rather than every integration being bespoke.
What they give you
- A common, discoverable interface for tools and capabilities, so integrations are written once rather than per agent
- A defined place for a gateway or proxy to sit, because protocol traffic can be intercepted, authenticated, logged and rate-limited like any other protocol
- Portability of tools and agents across frameworks and vendors, which is what keeps this layer from becoming another lock-in
- Structured metadata about what a tool does, which makes policy decisions expressible rather than guessed
What they do not do
- Neither specifies authentication, authorisation, quotas, tenancy or audit at the depth an enterprise needs; those are left to whatever governs the traffic
- Neither prevents an agent from being talked into calling a tool it should not; prompt-level attack surface is a security problem, not a protocol one
- Neither gives you spend control, since an agent’s call count is decided at runtime
- Neither is a component you deploy or a bill you receive, so they cannot be compared against the products below
The practical consequence for this comparison: if your organisation is heading towards many agents and many MCP servers, the question is which of these three products is positioned to be the enforcement point for that traffic, and that is a product decision even though MCP and A2A are not products.
LiteLLM

LiteLLM is a library that grew into a proxy, and that lineage explains almost everything about it. It began as a Python package that normalised a very long list of provider APIs to the OpenAI call shape; the proxy server wraps that translation layer in virtual keys, team budgets, spend tracking, caching, rate limits and an admin UI. Provider coverage is its main asset and it is a real one in a market where models and endpoints appear constantly. Self-hosted is its native mode: you run the proxy, a Postgres for keys and spend, and a Redis for shared limits and cache, and the ops burden is yours.
Pros
- The broadest provider and model coverage of the three, which is exactly what the translation layer in front of a fast-moving model market should be good at
- Genuinely self-host-first, with keys, budgets, spend tracking and the admin UI in the open-source proxy rather than gated behind a hosted plan
- Configuration is YAML in version control, so model lists, fallback chains and load-balanced deployment groups are reviewable artifacts rather than console state
- No vendor in the data path and no vendor to leave, which settles data-residency arguments before they start
Cons
- Python in the request path means the proxy tier needs deliberate concurrency and worker sizing under high request rates with many concurrent long-lived streams, and that tuning is your problem
- You operate three components and own the retention policy for a Postgres that fills with prompt and completion payloads
- Enterprise governance — deeper SSO, audit and guardrail integrations — sits behind a paid tier, so the open-source build is not the whole feature list
- Fast-moving surface area with frequent releases, which means pinning versions and reading release notes is ongoing work rather than a one-off
Best for: Teams with a named platform owner who want maximum provider coverage, full data control, and no vendor in the request path, and who accept that the gateway’s uptime is now their responsibility.
Pricing: Open source with no licence cost for the core proxy plus a paid enterprise tier for advanced governance and support; the real cost of the self-hosted path is infrastructure for three components plus the operator time to keep them healthy.
Portkey

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks, and generally available for enterprises. Get this right when you evaluate it, because it changes what you are buying. The product is a managed enterprise gateway — unified provider API, configurable routing and fallback, virtual keys, caching, guardrails, prompt management and request-level observability — and its centre of gravity has moved to security and governance under its new ownership. The gateway core has an open-source lineage and a lightweight, edge-deployable design, but the buying decision for the enterprise product is now partly a decision about a security vendor’s platform.
Pros
- Managed operation removes the three-tier stateful service from your team’s responsibility entirely, which is the correct trade for most product teams
- The governance surface enterprises actually get blocked on — SSO, role-based access, audit trails, budget enforcement, guardrail policy — is product rather than assembly
- Security and AI-runtime protection are the organising idea rather than a plugin, which is a genuine advantage if your security team owns AI risk
- Lightweight gateway design with an edge-deployable heritage, so the added hop can sit close to the caller rather than in one distant cluster
Cons
- You are now buying into a security vendor’s platform, with the roadmap, packaging and procurement posture that implies — which is exactly what some buyers wanted and exactly what others did not sign up for
- A managed gateway in the request path means prompts traverse a vendor’s infrastructure unless you deploy a mode that avoids it, which is disqualifying for air-gapped and strict-residency environments
- Enterprise motion rather than self-serve for the full product, so serious evaluation is a conversation rather than a container
- Governance depth you may not need: a team that only wanted routing and spend visibility is paying for a platform
Best for: Enterprises that want an AI gateway they do not operate, where security and governance ownership sits with a security team and a Palo Alto Networks relationship is an asset rather than a complication.
Pricing: Commercial, sold on enterprise agreements with usage-based metering on gateway traffic rather than a single public list price; the practical variable is how the governance and security capabilities are packaged for your requirements.
Kong AI Gateway

Kong AI Gateway is Kong Gateway plus a family of AI plugins, and the bet is that LLM traffic is another protocol on the gateway you already run. Its product framing now covers LLM, MCP and agent-to-agent traffic through the same gateway and the same control plane, which is the most explicit position of the three on where agent governance belongs. The AI plugins do provider proxying, load balancing across models, prompt guarding and templating, semantic caching, and request and response transformation, and every plugin you already use for authentication, authorisation, logging and rate limiting applies to LLM routes unchanged.
Pros
- One gateway of record, one policy layer, and one team that already knows the tooling — an organisational advantage before it is a technical one
- LLM, MCP and agent-to-agent traffic governed through the same control plane, which is the coherent answer if agent traffic becomes a service-mesh-shaped problem
- Existing authentication, authorisation, rate-limiting and logging plugins apply to model routes, so governance is not duplicated in a second system
- Deployable self-hosted or with a managed control plane, so the operating model is a choice rather than a constraint
Cons
- The advantage is almost entirely conditional on already running Kong; adopting it for LLM traffic alone is a lot of gateway for a narrow purpose
- The open-source and enterprise split across the AI plugin set is the first thing to check, because an evaluation on the open-source build may not represent what you would ship
- LLM-specific depth — provider coverage, prompt management, per-request cost attribution — trails the specialists whose entire product is model traffic
- Configuration lives in Kong’s model, so LLM policy is expressed in gateway terms rather than in model terms, which is a translation your AI engineers have to learn
Best for: Platform teams already running Kong who want one control plane and one policy layer covering REST, LLM, MCP and agent traffic rather than standing up a second LLM-only hop.
Pricing: Open-source gateway with no licence cost, plus an enterprise subscription for the advanced plugin set and the managed control plane; which AI plugins fall on which side of that line is the question that decides your actual cost.
How to choose
Do not start from the features. Start from three questions about your organisation, in this order, and stop at the first one that gives a clear answer.
Do you already run Kong? If yes, and your requirements are covered by the AI plugins available in your tier, choose Kong. The marginal operational cost is near zero and you avoid a second control plane. This single question resolves more of these evaluations than any feature comparison.
Do you have a named platform owner with capacity for a critical-path stateful service? If no, do not self-host. A managed gateway is the correct answer, and a gateway you cannot maintain is worse than one with fewer features. If yes, and data control or provider coverage matters, LiteLLM is the strongest self-hosted option.
Who owns AI risk in your organisation? If it is a security team with a platform relationship, Portkey as Prisma AIRS AI Gateway lines up with how the decision will actually be made and approved. If it is the application team, the security-platform framing is weight you do not need.
| LiteLLM | Portkey / Prisma AIRS | Kong AI Gateway | |
|---|---|---|---|
| Architectural bet | A library that became a proxy you run | A managed gateway inside a security platform | LLM traffic as another protocol on the gateway of record |
| Who operates it | Your platform team | The vendor | The team that already runs Kong |
| Data path | Entirely yours | Vendor infrastructure unless deployed otherwise | Yours, self-hosted or with a managed control plane |
| Strongest asset | Provider and model coverage | Governance and security as product | One control plane for REST, LLM, MCP and A2A |
| Biggest cost | Three components and their uptime are yours | Buying into a security vendor’s platform | Conditional on already running Kong |
| Air-gapped viable | Yes | No for the managed path | Yes, self-hosted |
| Picks itself when | You need data control and the widest provider list | You want the gateway operated for you and security owns AI risk | Kong is already your policy layer |
Then validate with a two-week trial that tests what a feature grid cannot.
- Shadow one real production workload through each candidate for a week, and look at p99 and time-to-first-token rather than averages.
- Force a provider failure and inspect what each does — how the fallback fires, what the client sees, whether retries multiply, whether streams truncate cleanly.
- Have the person who will actually answer the security review answer one from each product’s audit surface. Not a demo: a real question, like which keys called which model last Tuesday.
- Write down the exit cost for each candidate. For LiteLLM it is a config port. For a managed platform it is prompt templates, saved configs, and however much history you cannot export.
The wider category, including options none of these three cover well, is mapped in the AI gateway hub. If a specific one of these three is already in place and you are looking to leave, Portkey alternatives and LiteLLM alternatives segment by reason.
Frequently asked questions
Is Portkey still Portkey?
The product continues as Prisma AIRS AI Gateway, part of Palo Alto Networks, and is generally available for enterprises. What changes is the framing of the purchase: an independent AI-gateway startup and a gateway inside a large security vendor’s portfolio are different things to buy, with different roadmap dynamics and different procurement paths. Write that into your evaluation rather than around it — for a company that already buys from Palo Alto Networks it is a simplification, and for a company that deliberately chose a focused independent vendor it is a change worth re-examining.
Can I run more than one of these?
Yes, and it is more common than it sounds. A frequent shape is Kong as the gateway of record at the edge for policy, authentication and audit, with LiteLLM behind it doing provider translation and model routing. That gives you one governance surface and the widest provider coverage. The cost is two hops and two things to operate, so only do it if both jobs are genuinely needed.
Which one is fastest?
The wrong question. All three add a network hop, and on a multi-second streaming generation that hop is a rounding error. What actually matters for latency is whether the gateway streams rather than buffers, whether it reuses upstream connections, and whether it sits in the same region as the caller. Any of the three configured badly will be slower than any of the three configured well.
Does self-hosting LiteLLM save money versus a managed gateway?
Only if you count correctly. The licence cost is zero and the infrastructure cost is small. The real cost is the engineering time to operate a critical-path service with a database and a cache, plus the incident cost the first time it goes down and takes every AI feature with it. At small scale a managed gateway is usually cheaper in total. At large scale, or where data control is a requirement rather than a preference, self-hosting wins on more than price.
What if I only need routing and spend visibility?
Then all three are more product than you need, and the lighter end of the category is the right place to look. Buying a governance platform to act as a proxy is how teams end up paying for capabilities nobody enables.
Related reading
- Best AI gateways — the full category map and where each type of gateway fits.
- Best self-hosted LLM gateways — what the self-hosted path actually costs to operate.
- Portkey alternatives — segmented by the four real reasons teams leave.
- LiteLLM alternatives — segmented by reason, including not wanting to operate a proxy.
- Best AI gateways for enterprise — SSO, RBAC, audit and the procurement blockers that eliminate most options.
- Best open source AI gateways — licence models and what each project keeps behind an enterprise tier.