An enterprise AI gateway evaluation almost never fails on features. It fails in a security review, six weeks after engineering has already built against a product, on a question nobody asked first: where does the prompt go, who can see it afterwards, and what stops a team spending the department’s quarterly budget in an afternoon.
The pattern is predictable. Engineering picks on developer experience and provider coverage. Security asks about egress, redaction and audit. Procurement asks about deployment mode, residency, and what happens to the contract if the vendor is acquired. Then the shortlist changes and the six weeks are gone. So invert the order: establish the control list, eliminate everything that cannot satisfy it, and only then compare developer experience among the survivors. That usually leaves three or four candidates rather than fifteen.
There is also a newer item on the list that most enterprises have not answered. The gateway is no longer only about chat completions: MCP servers are being wired to internal systems by individual teams, agents are calling other agents, and none of that traffic passes through anything that logs or authorises it.
Key takeaways
- Six controls decide the shortlist: redaction before egress, an audit trail of who called which model with what data, hard budget enforcement, SSO and RBAC, model allow-lists, and an acceptable deployment mode.
- Hard enforcement is the control most often faked. A budget that emits an alert is a reporting feature; one that returns an error is a control.
- MCP and agent-to-agent traffic is the next ungoverned surface, and Kong’s product now covers LLM, MCP and A2A through the same gateway.
- Several independent vendors here are now inside incumbents: a bigger balance sheet, someone else’s roadmap.
The control list that decides the shortlist
Run these six before any demo. Each eliminates candidates, and each has a version that looks satisfied on a feature page and is not.
PII redaction before egress. Identifiers removed or tokenised before the request leaves your network boundary, not after it arrives somewhere. That distinction eliminates every managed gateway for genuinely sensitive workloads, because redaction inside the vendor’s service means the unredacted prompt already crossed the boundary to get there. Ask where the step executes, then what happens on a partial match — a redactor that catches a formatted account number and misses the same number with spaces provides assurance rather than protection. The detectors are their own category, LLM guardrails.
An audit trail of who called which model with what data. Four fields, and most products give two. Identity must be a real principal resolved through your identity provider, not an API key label somebody typed. Model version matters, because “which model saw this data” is what a regulator asks and a silently repointed alias makes it unanswerable. The data is the hard part: full prompt retention is the most useful record and the largest new store of sensitive data you have ever created, so the honest design is a hash plus a redacted copy plus a retention limit. And the log must be tamper-evident and exportable to your SIEM.
Budgets with hard enforcement. The most overstated control here. An alert at eighty percent is a reporting feature. Enforcement means the gateway returns an error at the limit, and the interesting questions are the edges: how the counter is shared across replicas, whether it is evaluated before or after the call completes, and what happens to a stream that crosses the limit mid-response. Streaming is where most implementations quietly fail open.
SSO and RBAC. SSO with group mapping, so access follows the same joiner-mover-leaver process as everything else. RBAC granular enough for the roles you have: who changes routing, who sees prompt contents in logs, who raises a budget, who adds a model to the allow-list. A product where every administrator can do all four is one where the audit trail has a single meaningful entry.
Model allow-lists. A default-deny list of which models each team may call, maintained centrally. It stops an engineer routing regulated data to whichever model topped a leaderboard last week, and it is how you enforce a data-processing agreement — the agreement covers specific providers, and a gateway that can reach any provider will reach one you have no agreement with.
A deployment mode your security team accepts. A short ladder: managed SaaS, single-tenant in the vendor’s cloud, your own VPC, on-premise, air-gapped. Each rung eliminates vendors, and the answer is structural rather than negotiable — a vendor whose product has never run air-gapped will say yes in a sales call and find out during implementation.
Needs first-hand data: For each shortlisted gateway, send one request containing three PII formats — a formatted account number, the same number with separators, and a name in free text — and capture where each was redacted using network traces rather than the vendor’s log. That tells you whether redaction happens before egress or after ingest.
Needs first-hand data: Set a small hard budget on a test project and exhaust it three ways: concurrent requests across replicas, one long streaming response that crosses the limit mid-stream, and a request that fails after the provider was billed. Record which of the three the gateway actually blocks.
MCP and A2A: the standards underneath, not products
Two specifications now shape what an enterprise gateway has to govern, and neither is something you buy. MCP standardises how a model gets tools and context from a server; A2A standardises how agents talk to each other. Both are wire protocols with no vendor, no dashboard and no bill, and their relevance is the traffic they create that your existing controls were not designed for.
The governance problem is simple once stated. An MCP server lets a model act on a real system — a database, a ticketing tool, an internal API — authorised by whatever the person who connected it configured, frequently a token with more scope than the task needs. Agent-to-agent traffic is worse: the caller is not a person, and the chain of delegation is exactly what an audit trail needs and exactly what nothing records.
What they give you
- A common shape for tool and agent interfaces, so a gateway can inspect and authorise this traffic generically rather than per integration
- The prospect of one policy plane over model, tool and inter-agent calls under a single identity model
- A tool server or agent written against the spec is reachable from more than one runtime
What they do not do
- Neither carries an enterprise authorisation model: identity, scope, consent and delegation are yours to impose
- Neither produces an audit record of who authorised an action, which is the field a regulator asks for
- Neither constrains what a tool server can do, so a permissive MCP server is a permissive path into production
- Adoption is uneven and moving, so a conformance claim tells you less than testing the pair you will run
The gateway is becoming the policy enforcement point for all of it
The argument is the one that put REST traffic behind an API gateway twenty years ago, and it lands harder here.
Model calls, tool calls and agent-to-agent calls are all outbound machine-to-machine requests carrying sensitive data, costing money, originating from code no central team reviews. Enforcing policy in each application means once per application, in whatever language it is written in, maintained by whoever has time. Enforcing it at a hop means once — including for the application nobody remembers deploying.
Kong’s product now covers LLM, MCP and agent-to-agent traffic through the same gateway, the clearest statement of that direction from an incumbent. Control plane, identity model and audit trail become shared across all three rather than three separate governance projects, and the one you have not started is usually the agent traffic.
The cost of concentrating this much in one hop should be said plainly: its failure stops model calls, tool calls and agent workflows simultaneously, and its misconfiguration is a security incident across all three at once. Enterprises answer that as they do for API gateways — replicas, config as code with review, staged policy rollout — provided somebody owns it as a job. If your agent architecture is still being designed, AI agent frameworks is where the traffic patterns come from.
Consolidation is now a procurement question
Several of the independent vendors in this category are no longer independent. Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks; Helicone has announced it is joining Mintlify. In adjacent categories the pattern repeats: Arize with Dynatrace, Traceloop joining ServiceNow, Galileo now part of Cisco, Promptfoo now part of OpenAI, Prompt Security now part of SentinelOne, OpenMeter now OpenMeter by Kong.
For an enterprise buyer this cuts both ways, and both are legitimate inputs rather than reasons to panic.
A startup gateway in the request path of regulated workloads was always a vendor-risk conversation, and a large parent balance sheet removes the going-concern question. Certifications, regional deployment options and support commitments usually improve under an incumbent.
Against that, the roadmap now serves a platform strategy that is not yours. A gateway inside a security portfolio will be developed as a security control — good if that is why you bought it, less good if you needed provider coverage or prompt tooling. Pricing power shifts too: a standalone product can be folded into a bundle, and the standalone SKU is not guaranteed to exist at your next renewal.
The response is contractual. Get support and roadmap commitments in writing, ask what the migration path is if the product is folded into a suite, and keep the integration shallow enough that leaving is a configuration change. That last part is engineering, and the only part you fully control.
Kong AI Gateway

Kong AI Gateway is plugins on Kong’s existing API gateway, which is why it wins where Kong already runs: the authentication model, audit pipeline, tenancy and change process are the ones security already approved. Its product covers LLM, MCP and agent-to-agent traffic through the same gateway, so the governance story spans all three rather than needing three programmes.
Pros
- Reuses an already-reviewed control plane for auth, rate limiting, logging and audit export
- One policy plane across REST, LLM, MCP and A2A under a single identity model
- Runs entirely in your own infrastructure, restricted and disconnected environments included
- Mature multi-workspace tenancy from years of API gateway work, which is what per-team governance needs
Cons
- Without an existing Kong footprint you are adopting an API gateway platform to get an LLM feature
- Configuration-first: application teams find it heavier than a managed hop
- Prompt lifecycle, evaluation and experiment tracking need separate tools
Best for: Platform teams already operating Kong who want one control plane across REST, LLM, MCP and agent traffic.
Pricing: Open-source core gateway with paid enterprise tiers priced by deployment scale and support; the AI plugins follow the same split.
Portkey

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks, generally available for enterprises. That repositioning is the relevant fact here: the gateway is presented as a security control for AI traffic rather than a developer convenience, which is the framing a security review wants. Underneath it is still a config-driven gateway with routing, fallback, caching, prompt management, budgets and tracing.
Pros
- Governance framing arrives pre-built for a security review, shortening the internal sell
- Broad control surface in one hop: routing, budgets, prompt versioning, tracing, guardrails
- A large security vendor’s compliance and support posture behind a component in your path
Cons
- The roadmap answers to a security platform strategy, so provider coverage and developer tooling are no longer first priority
- A managed hop means the request has already crossed your boundary before inspection happens
- Shallower than dedicated evaluation and prompt-engineering tools, which you will still buy
Best for: Enterprises wanting a managed gateway whose governance controls are already framed as security controls, and can accept a vendor-hosted hop.
Pricing: Enterprise contracts for the governance deployment with usage-based tiers underneath, not itemised publicly. Similar options are in Portkey alternatives.
TrueFoundry

TrueFoundry comes at this from the platform side: an AI gateway inside a broader deployment and model-serving platform, designed to run in your own Kubernetes cluster or VPC. Where the blocking requirement is that nothing leaves the network, that answers the egress question by design rather than by policy setting — and you are evaluating a platform rather than a proxy.
Pros
- Built for in-VPC and in-cluster deployment, so the egress objection is structural rather than configured
- Gateway sits alongside model serving, which suits running hosted and self-hosted models together
- Per-team access control, budgets and attribution built for a platform team serving many product teams
Cons
- Adopting a platform, so evaluation, rollout and training cost far more than a proxy’s
- Requires competent Kubernetes operations, a real constraint even in large organisations
- Smaller vendor than the incumbents, which procurement will treat as vendor risk
Best for: Central platform teams that must keep model traffic inside their own VPC and also serve self-hosted models.
Pricing: Enterprise contracts scaled by deployment size and support, with the software in your own infrastructure so compute is your cost rather than a metered charge.
Azure API Management

Azure API Management is a general-purpose API gateway that has grown AI-specific policies, and in Microsoft-centred enterprises it is the path of least resistance: identity through your existing directory, policy as code your platform team already reviews, and a gateway already fronting your other APIs. It rarely wins on being the best LLM gateway. It wins on being already procured, approved and in the network diagram.
Pros
- Already-approved procurement and security posture where Azure is the standard, removing most of the review cycle
- Access control integrates with the directory your joiner-mover-leaver process already drives
- Policy engine handles token-aware rate limiting, caching and backend load balancing as configuration, with topologies that include gateways inside your own network
Cons
- Provider-agnostic in principle, Azure-shaped in practice, so multi-cloud model access is more work
- Policy expression is verbose, so application teams will want an abstraction over it
- No prompt lifecycle or evaluation tooling; a traffic control plane only
Best for: Enterprises standardised on Azure that want AI traffic governed by the same gateway and policy pipeline as their APIs.
Pricing: Consumption and tier-based pricing on gateway capacity and request volume, billed within the cloud subscription rather than as a separate contract.
Apigee

Apigee is the same argument from the Google Cloud side: mature API management, with the quota, developer-portal and analytics machinery that comes from governing APIs at scale, now applied to model traffic. Its strongest card is the hybrid model, running the data plane in your environment with a managed control plane.
Pros
- Deep API governance: quotas, onboarding, versioning and analytics for many internal consumers
- Hybrid deployment keeps the data plane in your own environment
- Existing API governance processes and reviewers transfer directly to model traffic
Cons
- Enterprise-weight learning curve; overkill unless you govern many consumers
- Token accounting, streaming semantics and model fallback are less native than in purpose-built gateways
Best for: Large organisations already running Apigee that must expose model access to many internal teams with quotas.
Pricing: Enterprise subscription tiers based on API call volume and deployment topology, contracted through the cloud provider.
Amazon Bedrock

Bedrock is not a multi-provider gateway, and treating it as one causes most of the confusion in these evaluations. It is a managed model service many enterprises use as their gateway, because the controls arrive with it: authorisation through IAM, API call logging through the platform’s audit trail, private network access, and guardrail policies on requests and responses. For “governed model access inside our cloud boundary” it is a strong answer with nothing added.
Pros
- Governance comes from cloud primitives security already audits: identity policy, audit logging, private networking, key management
- No new vendor in the request path and no new contract, removing an entire procurement cycle
- Model access scoped by identity policy is a natural implementation of an allow-list
Cons
- Only the models on the platform, so no route to providers outside it and no capacity hedge
- Portability is poor by construction: request shape, identity model and guardrail config are platform-specific
- No cross-provider routing or fallback, which is why most teams wanted a gateway at all
Best for: Enterprises standardised on AWS needing governed access to a curated catalogue inside their existing cloud boundary.
Pricing: Consumption-based on tokens processed, with separate metering for guardrail evaluation and provisioned capacity, billed in the existing cloud account.
Envoy AI Gateway

Envoy AI Gateway puts provider credentials, model routing and token-aware rate limiting into an Envoy-based data plane, built in the Envoy ecosystem under open governance rather than by a single vendor. For a platform team that is a procurement argument, not an ideological one: a project with multi-vendor governance cannot be acquired out from under you.
Pros
- Open governance removes the single-vendor acquisition risk that now applies across this category
- The data plane is Envoy, already in your request path and already understood by your platform team
- Kubernetes-native config fits GitOps and policy-as-code review, and it runs entirely in your infrastructure, so the egress question is answered by architecture
Cons
- No commercial entity to hold to an SLA, which some procurement processes cannot accommodate
- Assumes genuine Envoy and Kubernetes competence; the hardest option here without a platform team
- Admin UI, prebuilt reports and prompt lifecycle are out of scope and must be assembled
Best for: Platform teams with real Envoy expertise who want LLM and agent policy in the data plane they already run.
Pricing: Open source with no licence cost; the cost is infrastructure plus the engineers who operate it.
LiteLLM

LiteLLM earns an enterprise shortlist place for one reason: it is the broadest self-hosted OpenAI-format translation layer available, so provider coverage is solved with no prompt leaving your network. It runs as a proxy with virtual keys, budgets and spend logs, which maps onto per-team governance. The caution is that it is open core, and the features an enterprise needs are the ones most likely to sit above the paid line — see open source AI gateways.
Pros
- Widest provider coverage of any self-hosted option, local and private endpoints included
- Virtual keys with budgets and spend logs give per-team attribution and enforcement without a vendor hop
- Nothing leaves your network, so redaction and inspection genuinely happen before egress
Cons
- SSO, granular RBAC, audit logging and organisation management are the controls most likely to sit in the paid tier, so verify each rather than assuming
- A stateful service with a database in the request path, with its upgrade and capacity path yours
- Governance ergonomics are thinner than the API management platforms: fewer reports, less mature change management
Best for: Platform teams needing broad provider coverage with zero egress, who can operate a stateful service themselves.
Pricing: Open source for the core plus a paid enterprise tier for organisation-level features; self-hosting moves the rest into infrastructure and staff time.
How to choose
Work top to bottom. Each step removes candidates, and this order saves the six weeks above.
One: fix the deployment mode. Managed SaaS, single-tenant, in-VPC, on-premise or air-gapped. Usually a policy decision rather than a technical one, and the single largest eliminator: everything managed disappears the moment the answer is “nothing leaves the network”.
Two: check hard enforcement, not the feature list. Set a small budget in a trial and try to break it three ways. Products that only alert reveal themselves here.
Three: check where redaction runs. Not whether it exists — where it executes, verified with a network trace rather than a vendor log. Redaction after ingest does not satisfy a before-egress requirement, however good the detectors are.
Four: test the audit record against a real question. Take an actual regulatory question — “list every request in March where customer data reached a third-party model, and who authorised it” — and answer it from the product’s logs. Most fail on identity or model version. Only then compare developer experience: among survivors those differences are recoverable, and the controls are not.
| Gateway | Deployment shape | Picks itself when |
|---|---|---|
| Kong AI Gateway | Self-hosted, any environment | Kong already fronts your traffic |
| Portkey / Prisma AIRS | Managed | You want governance without operating anything |
| TrueFoundry | In your VPC or cluster | Nothing may leave the network and you serve your own models |
| Azure API Management | Managed with in-network options | You are standardised on Azure |
| Apigee | Hybrid, data plane in your environment | Many internal teams need self-service with quotas |
| Amazon Bedrock | Managed inside your cloud account | Governed access to one curated catalogue is enough |
| Envoy AI Gateway | Self-hosted data plane | You run Envoy and want no single-vendor risk |
| LiteLLM | Self-hosted proxy | Provider breadth with zero egress is the requirement |
If compliance is the driving constraint rather than one requirement among several, AI gateways for compliance goes deeper on residency and evidence, and self-hosted LLM gateways is where the operational cost of running the hop is quantified. The architectural background is in the AI gateways hub, and the head-to-head is LiteLLM vs Portkey vs Kong.
Frequently asked questions
Can we use our existing API gateway instead of a purpose-built AI gateway?
Often yes, and it is the fastest route through security review, because the identity model, audit pipeline and change process are already approved. What you give up is LLM-native behaviour: token-aware limits, streaming semantics, cross-provider fallback and per-model cost attribution are absent or implemented as policy. If governance is the goal and provider flexibility secondary, the existing gateway is the pragmatic answer.
Does PII redaction at the gateway satisfy a data-protection requirement?
Only if it executes before data leaves your boundary, and only if you have tested the detectors against your own formats. Redaction inside a vendor’s managed service means the unredacted prompt already crossed the boundary. Treat it as one layer, not as the control that makes an external model safe for regulated data.
How do we govern MCP servers teams have already connected?
Inventory first — you cannot govern what you cannot list. Then require MCP traffic through the same gateway as model traffic so it inherits identity, authorisation and audit, and review the credential scope each server holds. The common finding is a token with far more permission than the tool needs.
What is the biggest mistake in an enterprise AI gateway evaluation?
Letting engineering pick first. Establish the control list before the demo, not after the integration.
Related reading
- Best AI gateways — the architectural decision behind every product here.
- Best AI gateways for compliance — residency, evidence and audit as the primary constraint.
- Best self-hosted LLM gateways — what running the hop yourself costs in operations.
- Best open source AI gateways — where the open-core line sits on enterprise features.
- Best LLM guardrails tools — the detectors behind redaction and injection screening.