Buyer’s Guide

Best AI Gateways for Compliance and Data Residency

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • compliance
  • ai-gateway
  • security
  • infrastructure

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

AI compliance reviews rarely fail on a certification. They fail on five questions the team cannot answer: where did this prompt go, what was in it, who sent it, which model saw it, and how long is it kept. Nobody decided any of it. It was all decided months earlier, by accident, by whoever wrote the first integration and whoever set the default log level.

That is the shape of the problem. Every answer an auditor wants is a design decision that already happened. The application calls a provider SDK directly, so “where did it go” is whatever was in an environment variable. The logger writes request bodies at debug level, so “what was in it” sits in a log index with a retention policy set for web traffic. The API key is shared across three services, so “who sent it” is unanswerable in principle. None of that was negligence — it was the default, and the default is what you are being audited on.

An AI gateway earns its place because it makes those five questions answerable in one place instead of unanswerable in twenty. That is the value proposition, well ahead of routing or caching. What it does not do is make you compliant, and no product here can: certifications belong to vendors and terms belong to contracts, both of which you must verify yourself against current documents. The architecture decides whether the answers exist at all.

Key takeaways

  • Five questions decide an AI compliance review: where the prompt went, what was in it, who sent it, which model saw it, and how long it is kept. Each is a design decision, usually made by default.
  • Retention is three separate policies — the provider’s, your gateway’s, and your observability platform’s. A zero-retention arrangement upstream means nothing if your own logs keep full prompts for ninety days.
  • Regional endpoints narrow the surface, but the only structural residency guarantee is data that never leaves infrastructure you control. For frontier models that makes residency a contractual question, not a technical one.
  • Redaction before egress is the control with the most effect per hour spent, and it is lossy in both directions: entity recognition misses things, and removing context sometimes removes the answer.
  • An audit trail is not a debugging trace. Different consumers, different retention, and conflating them is how teams keep raw prompts far longer than they meant to.

The five questions, and where each answer has to live

Where did this prompt go? Not “which vendor” — which endpoint, in which region, through which intermediaries. A hosted gateway is an intermediary. So is an observability platform that receives prompt content, and so is any evaluation service you send production traffic to. Draw the actual graph of systems that see prompt text; it is always larger than people expect, and the extra nodes are usually tooling nobody classified as data processing.

What was in it? This requires knowing your own request composition. A prompt is a system message, retrieved documents, conversation history and user input, and the risky content is usually retrieval, not the user. A retrieval layer that pulls from an internal store with mixed sensitivity will put things in a prompt that no one intended to send anywhere.

Who sent it? Attribution needs identity to survive from the end user to the model call. Shared API keys destroy it. The fix is per-service and ideally per-tenant credentials issued by the gateway, with the calling user propagated as a request attribute — which is a change to how your applications call the gateway, not a gateway setting.

Which model saw it? Model identity means version, not family. Answering “which model produced this output” eight weeks later means recording the resolved model version per request, because a provider alias moves and your logs will otherwise name a model that was not the one running.

How long is it kept? Three separate answers, which is the next section.

Four of the five are properties of the request path, and a single request path is the only place you can enforce them. That is the whole argument for a gateway.

Retention is three decisions, not one

Teams collapse retention into one question and get it wrong in a specific, repeatable way.

What the provider keeps. Vendor terms differ, and the shapes to look for are: a default retention window for abuse monitoring, a zero-retention or no-training arrangement available on request or by tier, and exemptions where content is retained despite that arrangement if it trips a safety review. Those are policy shapes, not facts about any particular vendor — read the current terms and the data processing addendum for the vendor and tier you are actually on, and get the arrangement in the contract rather than in a support ticket. A zero-retention promise you cannot point to in an executed document is not a control.

What your gateway logs. This is the one that undoes the rest. A gateway that writes full request and response bodies to Postgres for ninety days means the provider’s retention window is irrelevant: your own copy is the longest-lived one, it is in your primary database, and it is probably not encrypted differently from anything else there. Decide deliberately between logging metadata only, logging redacted bodies, logging full bodies for a short window, or sampling. Then verify it, because “we log metadata” and “we log metadata plus the request body on error” are different systems and the second is what most implementations actually do.

What your observability platform stores. Traces containing prompt and completion text inherit the trace store’s retention, which was set for latency debugging and is often a year. This is where prompt content lives longest in most organisations, and nobody chose it. The LLM observability roundup covers which platforms let you drop or redact content at the collector, before it lands.

The composite retention of your system is the maximum of the three, not the minimum. Write the three numbers next to each other; the exercise takes an hour and usually finds a surprise.

Needs first-hand data: For one production request, trace its content through every system that stores it — gateway logs, application logs, trace store, evaluation datasets, cache, and any managed provider logging — and record the retention period and encryption posture of each. Publish the maximum, not the intended value. That single table is the most useful compliance artefact you can produce in a day.

Data residency: what is structural and what is contractual

Residency requirements come in three strengths and teams frequently satisfy the wrong one.

Regional endpoints. Most large providers offer endpoints in specific regions, and using one reduces the surface: inference happens in that region. What it does not tell you by itself is where request logs go, where abuse-monitoring copies go, or which support jurisdiction can access them. Those live in the terms, and they are the part that matters for a strict reading. Regional inference is necessary and not sufficient.

Self-hosted gateway plus a regional provider endpoint under contract. The gateway, the logs, the redaction and the audit trail stay in your infrastructure and your region, and only the redacted inference call crosses to a contractually bounded regional endpoint. This is the realistic architecture for most regulated teams, and the honest description of it is: technically controlled up to the boundary, contractually controlled beyond it.

Self-hosted inference. The only structural guarantee. Data never leaves infrastructure you control, so residency is a fact about your network rather than a clause. It is also, for frontier models, unavailable — you cannot run the largest closed models on your own hardware at any price. What you can do is run strong open-weight models yourself for the traffic that carries regulated data, and route the rest to a hosted provider. That split is a real architecture and it is what a self-hosted gateway is for; the operational side of it is in the self-hosted LLM gateway guide.

Be honest in the design document: unless you self-host inference, residency is a contractual question with a technical perimeter around it. Teams that write “data stays in the EU” without naming the clause that guarantees it are describing an intention.

One frequently missed corner: embeddings. A vector is derived from the source text and, with enough access, information about that text is recoverable from it. Treat vectors as carrying the sensitivity of what produced them and keep the store inside the same residency boundary as the documents.

Redaction before egress, and why it is lossy in both directions

Redaction at the gateway does more per hour of engineering than any other control here, because it is the only one that reduces what leaves your perimeter rather than recording that it left. Do it there rather than in each application, and it happens once, uniformly, on a path nobody can bypass by shipping a new service.

The mechanisms, in increasing order of what they cost you:

Pattern matching for structured identifiers — card numbers, national insurance numbers, account references. High precision, catches only what has a shape.

Named entity recognition for names, addresses, organisations, dates. This is where the honest limitations live. Recognition is a model, models miss things, and the misses are not random: unusual names, transliterations and non-Western formats fail more often, which means your redaction quality varies by customer demographic. That is worth knowing before you describe the control as complete.

Tokenised replacement — substitute a stable placeholder, and rehydrate the real value in the response on the way back. Considerably better than deletion because the model can reason about “CUSTOMER_1” consistently, and it needs a mapping store, which is now itself a store of sensitive data with its own retention question.

Then the tension nobody puts in the brochure: a redacted prompt sometimes loses the context that made the answer useful. Redact the address in a delivery query and the model cannot reason about the region. Strip identifiers from a complaint summary and you get a generic answer. The engineering work is not turning redaction on, it is deciding per use case what may leave — a product conversation with whoever owns the workflow.

My position: tokenise rather than delete wherever a stable reference will do, run the recognisers fail-closed on the small number of high-risk flows, and measure the miss rate rather than asserting it. The products that specialise in detection are in the guardrails guide.

Needs first-hand data: Assemble 200 real prompts, hand-label every entity that must not leave your perimeter, then run them through your gateway’s redaction. Record recall by entity type, and record it separately for non-Western names and formats. Also measure answer quality on the redacted prompts against the originals, because the recall number alone will make you over-redact.

The audit trail is not your debugging trace

These get conflated constantly, and the conflation is how organisations end up storing raw prompts for a year.

A debugging trace exists for engineers, contains full prompt and completion content because that is the point, is queried within days of the event, and should have short retention. A compliance audit trail exists for auditors and incident response, needs to be immutable, attributable to a specific user and service, and queryable long after the request — but does not need the prompt text at all in most regimes. Its fields are: timestamp, authenticated identity, service, tenant, resolved model and version, provider endpoint and region, token counts, policy decisions applied (redaction fired, guardrail blocked, budget rejected), and a content hash.

Notice what that list makes possible: you can prove which user’s request reached which model in which region eight weeks ago, and whether redaction ran, without retaining a single prompt. That is what lets you keep audit records for years and traces for two weeks. Store them in different systems with different retention and access control, and make the audit store append-only.

The pattern to avoid is one log stream serving both, retained for the longer requirement. That is the default outcome, and it is how a team meaning to keep prompts for a fortnight discovers it has kept them for a year.

Kong AI Gateway

Kong AI Gateway homepage

Kong AI Gateway extends a mature API gateway with LLM-specific plugins, and per its own product page it governs LLM, MCP and agent-to-agent traffic through the same gateway. For a compliance conversation that lineage is the point: authentication, mTLS, RBAC, audit logging and deployment topology are solved problems in Kong that you inherit rather than rebuild, and self-managed deployment means the request path and its logs stay in your infrastructure.

Pros

  • Self-managed deployment keeps prompts, logs and policy inside a perimeter you define, including air-gapped patterns
  • Inherits enterprise gateway plumbing — identity, mTLS, RBAC, audit logging — instead of reinventing it for LLM traffic
  • One control plane for REST, LLM, MCP and agent traffic, so policy is reviewed in one place rather than three

Cons

  • It is a gateway platform: real operational weight, and a team that has never run Kong is adopting two things at once
  • LLM-specific depth trails purpose-built AI gateways in areas like evaluation and prompt management
  • The enterprise features most compliance programmes need sit above the open-source tier, so budget accordingly

Best for: Platform teams already running Kong who want LLM, MCP and agent traffic under the same governed control plane as their REST APIs.

Pricing: Open-source core with enterprise subscription tiers for the governance, RBAC and support features, licensed per deployment rather than per token.

Portkey / Prisma AIRS AI Gateway

Portkey homepage

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks. Functionally it is a gateway with routing, caching, guardrails, budgets and observability under one config, and it can run hosted or inside your own environment. The corporate position is the relevant change for this article: your counterparty is now a security vendor, which shortens the vendor-review conversation and lengthens the procurement one.

Pros

  • Guardrails, redaction hooks, routing and per-key budgets in one control plane, so policy is not spread across services
  • Deployable into your own environment when the hosted data path is not acceptable
  • Request-level metadata and virtual keys give per-tenant attribution without shared credentials, and a security vendor as counterparty eases the risk review

Cons

  • Hosted mode puts a third party in the prompt path, which reopens exactly the question a compliance gateway exists to close
  • Consolidation under a large security vendor shifts pricing power and roadmap priorities away from self-serve users
  • Its declarative config is expressive enough to become a thing you must review and version like code

Best for: Enterprises that want routing, guardrails and governance in one product and prefer a security vendor as the contracting party.

Pricing: Tiered subscription with usage-based metering on requests, and enterprise agreements for self-managed deployment; enterprise pricing is not publicly itemised.

LiteLLM

LiteLLM homepage

LiteLLM’s proxy is the quickest way to get a self-hosted gateway in front of every model call. The decisive compliance property is simple: you run it, so prompts, logs, keys and policy never leave your network unless you configure them to. Virtual keys give per-service and per-tenant attribution, and its callback hooks are where redaction and audit logging get inserted.

Pros

  • Self-hosted by default, so the request path and every log are inside your own perimeter
  • Virtual keys per service or tenant replace shared credentials, which is what makes “who sent it” answerable
  • Callback hooks give you a place to enforce redaction and write an append-only audit record before egress
  • Open source, so a reviewer can read what it does with request bodies rather than trust a claim

Cons

  • Logging defaults deserve scrutiny: verify what lands in your database and for how long rather than assuming metadata only
  • You operate it on the hot path for every model call, with its own scaling and availability story
  • Compliance features are primitives you assemble, not a governed control plane with an audit UI a reviewer can be shown

Best for: Teams that need a self-hosted gateway with per-tenant keys and their own redaction and audit hooks, and have a platform owner for it.

Pricing: Open source with no licence cost plus a commercial enterprise tier; the real cost is infrastructure and the engineer who owns it.

Envoy AI Gateway

Envoy AI Gateway homepage

Envoy AI Gateway builds LLM routing and policy on top of Envoy and Kubernetes Gateway API, which makes it the natural choice for platform teams whose whole network already speaks Envoy. Everything runs in your cluster, policy is Kubernetes resources under GitOps review, and the observability and mTLS story is the one your service mesh already has — meaning your AI traffic stops being an exception to your existing controls.

Pros

  • Runs entirely in your cluster, so residency is a property of where the cluster is rather than a contract clause
  • Policy as Kubernetes resources means change review, version history and rollback come from Git rather than a vendor UI
  • Inherits Envoy’s mTLS, identity and telemetry, so LLM traffic is governed like every other service, with no single-vendor roadmap risk in the control plane

Cons

  • Requires real Kubernetes and Envoy competence; without it this is the hardest option here to operate
  • Younger than the general-purpose gateways, with a thinner set of LLM-specific features like prompt management and evaluation
  • No hosted option, so there is no low-effort path to try it

Best for: Kubernetes platform teams already standardised on Envoy who want AI traffic under the same in-cluster policy and GitOps review as everything else.

Pricing: Open source with no licence cost; you pay for cluster capacity and the platform engineers who run it.

Azure API Management

Azure API Management homepage

Azure API Management has become the default AI gateway for Azure-centric organisations, with policies aimed specifically at LLM traffic: token-based rate limiting, semantic caching, load balancing across model deployments, and emitting token metrics. The compliance argument is that it sits inside a subscription you already govern, so identity, network isolation, private endpoints, regional placement and diagnostic settings are the ones your platform team already configured.

Pros

  • Lives inside your existing Azure governance: identity, private networking, regional placement and diagnostics are already policy
  • Token-aware rate limiting and metrics per subscription key give per-consumer attribution without new plumbing
  • One API management layer for AI and non-AI traffic, so reviewers see one control plane

Cons

  • Azure-centric by construction; using it as a neutral multi-cloud gateway means fighting the gravity
  • Policy expressions are their own language and non-trivial policies become a maintenance burden
  • Diagnostic logging can capture more request content than you intend — verify what is written and where it is retained

Best for: Azure-standardised enterprises that want AI traffic governed by the same API management and network controls as everything else.

Pricing: Tier-based on gateway capacity units rather than per token, with the AI-specific policies available on the tiers that support them.

Amazon Bedrock

Amazon Bedrock homepage

Bedrock is a managed model platform rather than a gateway, and it earns its place because for AWS-standardised organisations it collapses several of the five questions into controls that already exist. Model access is IAM. The network path can be a VPC endpoint. Invocation logging goes to your own bucket under your retention and encryption. Regional model availability is explicit, so residency is a deployment fact rather than a clause you interpret.

Pros

  • IAM for model access means authorisation and attribution use the identity system your auditors already understand
  • Private connectivity via VPC endpoints keeps inference traffic off the public internet
  • Invocation logging writes to storage you own, so retention, encryption and access are your settings
  • Regional model availability is explicit, which makes a residency statement checkable rather than interpretive

Cons

  • Model catalogue is what the platform offers in your region, so a model you want may be unavailable where you must run
  • Single-cloud by construction; a multi-provider strategy still needs a gateway in front of it
  • Invocation logging captures full prompts and completions by default in many setups — that is a retention decision to make deliberately, not a feature to switch on

Best for: AWS-standardised teams that want model access governed by IAM, private networking and their own log retention rather than a third-party gateway’s.

Pricing: Usage-based per token by model, with provisioned throughput available as a committed capacity option; logging and storage are billed as the underlying AWS services.

Amazon Bedrock Guardrails

Amazon Bedrock Guardrails homepage

Bedrock Guardrails is the policy layer beside it: content filters with configurable strength, denied topics, word filters, contextual grounding checks, and sensitive-information policies that can block or mask entities on the way through. That masking capability is the one that matters for this article, because it puts redaction on the request path inside the platform rather than in each application, and it can be applied to models outside Bedrock too.

Pros

  • Entity blocking and masking put redaction on the request path as configuration rather than application code
  • Denied topics and grounding checks address hallucination and scope controls an auditor will ask about
  • Applies independently of the model, including to models not hosted in the platform, and policies are versioned resources so a change to a control has a history

Cons

  • Entity recognition misses cases, and the misses skew by name and format origin — measure recall rather than assuming it
  • Aggressive masking removes context the model needed, so quality regressions show up in exactly the workflows that carry sensitive data
  • Adds latency and per-request cost on both directions of every guarded call

Best for: Teams on AWS that need evidenced redaction and content policy applied uniformly, including to models outside the platform.

Pricing: Usage-based per unit of text evaluated per policy, metered separately from inference, so enabling more policies multiplies the per-request cost.

Google Gemini Enterprise Agent Platform

Gemini Enterprise Agent Platform homepage

Google’s enterprise AI platform, formerly Vertex AI, is now the Gemini Enterprise Agent Platform. Like Bedrock it is a managed platform rather than a gateway, and the compliance-relevant parts are the same in shape: IAM-based access, VPC Service Controls to build a perimeter around the services, customer-managed encryption keys, regional endpoints, and explicit controls over where data is processed and stored.

Pros

  • VPC Service Controls let you draw a service perimeter, which is a stronger network story than endpoint choice alone
  • Customer-managed encryption keys mean you hold the key material for data at rest
  • Regional endpoint and processing controls make a residency claim something you can point at in configuration

Cons

  • Single-cloud, and the platform rename means much existing documentation and internal runbooks use the previous name
  • Perimeter configuration is genuinely intricate; a misconfigured service perimeter is a common source of confusing failures
  • Platform-level logging can capture prompt content, so verify what is retained and where before assuming a default is safe

Best for: Google Cloud organisations that need a service perimeter, customer-managed keys and regional processing controls around model and agent traffic.

Pricing: Usage-based per token by model with separate metering for platform services such as agent runtime and evaluation; enterprise commitments available.

How to choose

Work from constraints, because in this category constraints eliminate most of the market before features matter.

One: decide whether prompt content may transit a third party. If not, every hosted gateway is gone and you are choosing between self-hosted options. This question decides more shortlists than anything else, and it needs a real answer from whoever owns the risk.

Two: decide whether residency must be structural or may be contractual. Structural means self-hosted inference for the regulated traffic, so a split architecture and open-weight models on that path. Contractual means a self-hosted gateway plus a regional provider endpoint, with the clause named in your design document.

Three: separate your audit store from your trace store now, before either has data in it. Retrofitting after a year of combined logs is a deletion project nobody wants.

Four: pick the gateway your platform team can operate. A governed control plane nobody maintains fails an audit more embarrassingly than a simple proxy that gets patched. If your team runs Kong, run Kong; if it runs Envoy, run Envoy AI Gateway; if neither, a hosted enterprise gateway under contract may be the lower-risk choice.

Five: verify certifications and terms yourself, from current documents. Nothing here is evidence about any vendor’s certifications, subprocessors or retention terms. Ask for the current report and the current data processing addendum, and read the exemptions.

OptionWhere the request path runsStrongest compliance property
Kong AI GatewayYours, self-managedMature enterprise gateway controls extended to LLM, MCP and agent traffic
Portkey / Prisma AIRSHosted or yoursOne governed control plane with a security vendor as counterparty
LiteLLMYoursSelf-hosted with per-tenant virtual keys and your own redaction hooks
Envoy AI GatewayYours, in-clusterPolicy as Kubernetes resources under GitOps review
Azure API ManagementYours, in-subscriptionAI traffic inherits existing Azure identity and network governance
Amazon BedrockManaged, in your accountIAM authorisation, private connectivity, logs in your own storage
Bedrock GuardrailsManaged, in your accountRedaction and masking applied uniformly on the request path
Gemini Enterprise Agent PlatformManaged, in your projectService perimeter plus customer-managed encryption keys

For the broader procurement view, the enterprise gateway guide covers RBAC and SSO depth, and the AI gateway hub covers the category as a whole.

Frequently asked questions

Does a zero-retention agreement with a provider mean prompts are not stored anywhere?

No. It constrains one party. Your gateway, your application logs, your trace store, your cache and any evaluation dataset built from production traffic are all separate stores with their own retention, and in most organisations one of them keeps prompt content far longer than the provider would have. Audit your own systems first — the upstream agreement is usually not the binding constraint.

Can I claim data residency if I use a provider’s regional endpoint?

You can claim inference happened in that region. Whether logs, abuse-monitoring copies and support access also stay in region is a matter of the terms, not the endpoint, and the honest phrasing in a design document names the clause that guarantees it. If the requirement will not tolerate a contractual answer, the only structural option is inference on infrastructure you control.

Should redaction happen in the application or the gateway?

The gateway, for the same reason authentication belongs there: one implementation on a path nobody can bypass. Applications drift, and the twelfth service will be written by someone who did not know the rule existed.

How long should we keep prompts and completions?

Shorter than you currently do, and separately from your audit trail. Most debugging value expires in days, so a short window for full content plus an indefinite append-only audit record of metadata and hashes gives you both incident response and evidence, without a store of raw prompts you would rather not have during discovery.