Buyer’s Guide

Best AI Gateways for Startups

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • ai-gateway
  • llm
  • infrastructure
  • cost

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

The pitch every AI gateway makes to a startup is that adoption costs one line. Change the base URL, keep your existing OpenAI client, and you inherit routing, fallback, caching, budgets and logs. It is an unusually honest pitch, because that really is how the integration works.

What it leaves out is the part that decides whether you should. A gateway is a dependency in the path of your product’s most important calls, and at seed stage a dependency is not free just because integrating it was. You are adding a component that can be down while your own code is fine, and a second place where model behaviour can change without you shipping anything.

The teams I have seen regret this went one of two ways. Some adopted a gateway in month one for a multi-provider future that never arrived, and carried an extra hop for a model string they never changed. Others waited until three services each had their own provider key, their own retry logic and their own idea of what a timeout means, and then found that retrofitting a gateway is not a base-URL change but an audit. So this is about the timing, and about what the base-URL swap genuinely buys before you rely on it.

Key takeaways

  • One product, one provider, one model: the OpenAI SDK plus a retry wrapper is less code than operating a gateway, and adopting one earlier buys a benefit you do not yet have.
  • Four things change the answer: a second provider, a per-customer cost question you cannot answer, a prompt you want to change without a deploy, and a key that now lives in more than one service.
  • A base-URL swap moves the transport, not the semantics. Provider-specific parameters, native tool-call shapes and streaming edge cases do not survive it intact.
  • Measure time to first routed request, not time to first request. Getting a call through is easy; getting a verified fallback that actually fires is the real setup cost.

Time to first routed request is the metric that matters early

Vendors advertise time to first request, which is the wrong number — any of these passes a curl in five minutes. What predicts whether the gateway earns its place is time to first routed request: a call that fails at the primary and comes back correctly from the secondary, with the failure in a log and the cost attributed to the right key.

That involves four things, and only the first is fast.

Getting a response through the gateway. Base URL, key, done. Minutes.

Getting the same behaviour you had before. Your existing prompt, parameters and tool definitions, with an identical output shape. This is where the first afternoon goes, and where the leaks below show up.

Getting a fallback that fires. Configuring a secondary is quick; verifying it is not, because you have to force the primary to fail. Most teams configure fallback, never test it, and learn during a real provider incident that the secondary rejected their tool definitions.

Getting attribution you trust. A dashboard showing spend is not attribution. Attribution answers “which customer caused this” without a spreadsheet, which means caller identity travels with every request — a header, a virtual key per tenant, or metadata your code sets deliberately.

Needs first-hand data: For each shortlisted gateway, time four checkpoints from a clean account: first response through the hop, byte-identical behaviour against your existing prompt suite, a forced primary failure that fell back correctly, and per-tenant cost visible for a synthetic customer. Publish the four numbers, not the first one.

You may not need one yet

This is the section the vendors will not write, so here it is.

If you have one product calling one provider with one model, the OpenAI SDK plus a retry wrapper is less code than operating a gateway. Not less code than integrating one — less code than living with one. The wrapper is maybe forty lines: exponential backoff with jitter, a timeout, a token-count log line, and a single function every call site goes through. You can read all of it. It has no uptime, no dashboard, no vendor and no configuration surface someone can change without a pull request.

Against that, a gateway at this stage is a dependency you carry for a benefit you do not have. Multi-provider routing matters when you have a second provider. Per-customer budgets matter when customers’ usage differs enough to care. Prompt versioning matters when someone other than the author changes prompts. None of that is true in month one, and a component in the request path that solves no current problem is pure risk.

Four triggers change the answer. Each is concrete enough to check today.

You have a second provider. Not “we might add one” — an actual second provider in code, because one model is better at extraction and another is better at summarising, or because a customer’s contract requires a specific vendor. The moment two providers exist, retry logic, cost accounting and prompt handling exist twice, and they will diverge.

A per-customer cost question you cannot answer. Someone asks whether your largest account is profitable, and answering means reconstructing spend from invoices that know only about API keys. If that is a project rather than a query, you have crossed the line. The cost tracking tools category attacks the same problem from the reporting side.

A prompt you want to change without a deploy. Usually the real trigger and rarely the stated one. When prompt changes need a deploy, iteration is bounded by your release process and a non-engineer cannot participate. A gateway with prompt storage decouples them — as do dedicated prompt management platforms, the honest alternative if that is your only reason.

A provider key in more than one service. When the second service needs a key you either copy the provider key — now in two environments, two CI configs, two laptops — or put something in front that issues scoped, revocable keys. Sharpest security consequence, and the trigger teams notice last.

If none of these is true, ship the product. Put the model call behind one function so adopting a gateway later is a one-file change, and go back to work.

Needs first-hand data: Count the call sites in your codebase that construct a provider client or read a provider API key. If the answer is one, you do not need a gateway yet. If it is more than three, you already needed one and the migration is now an audit rather than a config change.

The OpenAI API format: the standard underneath, not a product

Every “one line to adopt” claim in this category rests on one thing: OpenAI’s HTTP interface became the shape everyone else implemented. A messages array, a model string, /v1/chat/completions, SSE chunks for streaming. Providers, local runtimes and every gateway here speak it or expose a translation of it.

Be precise about what that is. Not a specification with a governing body, a conformance suite, or a version you can name in a contract — a widely copied interface that emerged because everyone wanted to be reachable by code already written against OpenAI. No vendor, no dashboard, no bill, and nobody guaranteeing that two compatible endpoints behave the same way.

What it gives you

  • A base URL and an API key are the whole integration surface, so adoption needs no application logic change
  • The official OpenAI SDK in any language becomes a client for any compatible endpoint, including a model running on your laptop
  • Your existing tests, load harness and request replay tooling keep working against the new endpoint
  • It makes a fallback provider technically reachable, which is the precondition for any routing at all

What it does not do

  • It does not cover provider-specific parameters, so cache control, reasoning-effort settings and provider-specific sampling knobs are passed as untyped extras or silently dropped
  • Compatibility is claimed, never certified: endpoints differ on tool-call formatting, streaming chunk boundaries, finish_reason values, token accounting and error bodies
  • It says nothing about behaviour. The same prompt to a different model is a different product, and no wire format makes that equivalent
  • It carries no tenancy, budget, audit or caching semantics, which is exactly why the gateways below are not interchangeable

A deeper treatment of which proxies stay closest to the format is in OpenAI-compatible proxies.

Where the base-URL swap leaks

The swap works. Then you hit the edges, usually in this order.

Provider-specific parameters. The common shape covers messages, temperature, max tokens, tools and streaming. It does not cover the parameters that were the reason you chose that provider: prompt-caching directives, reasoning-effort controls, safety settings, provider-specific stop and penalty behaviour. These live in an extras field if the gateway has one and are dropped if it does not. Silent dropping is the dangerous case — calls still succeed, they just stop using the cache you were relying on to control cost.

Native tool-call shapes. This is what breaks the fallback you never tested. Tool calling is where compatibility is thinnest: providers differ on whether arguments arrive as a JSON string or an object, whether multiple calls can return in one response, how a parallel call is represented, and what happens when arguments fail your schema. Your parser was written against the primary’s shape. The day fallback fires, the secondary’s shape arrives and your parser throws, in the middle of an incident that was already someone else’s fault.

Streaming edge cases. Streaming is where the semantics are least standardised and the failure is most visible. Chunk boundaries differ, so code that assumes a chunk is a token breaks. Some endpoints emit a final usage chunk and some do not, so token accounting silently zeroes. Mid-stream errors are the worst case: text is already rendered, and there is no clean way to represent “and then it failed”. And any layer that inspects the response — an output guardrail, a content filter, a naive HTTP client — has to accumulate it, turning token-by-token streaming into one delayed blob while every functional test still passes.

Token counting and cost. Providers tokenise and report usage differently, so a gateway’s cost figure comes from a price table it maintains. Compare its reported spend against the provider invoice once, early, before you build a customer-facing usage metric on top of it.

Rate limits move. Your limit was per provider key. Behind an aggregator you inherit a shared pool and can be throttled for reasons unrelated to your traffic. Behind your own gateway, its concurrency settings become the real limit, and you now have two places to tune.

Needs first-hand data: Build a small conformance suite — a tool call with two functions, a call with parallel tool calls, a streaming response, a mid-stream cancellation, and a request with one provider-specific parameter set — and run it against your primary and your intended fallback through the gateway. The diff is the actual cost of your fallback plan.

OpenRouter

OpenRouter homepage

OpenRouter is one OpenAI-compatible endpoint in front of a very large model catalogue across many providers, billed to one prepaid account. For a startup the value is mostly commercial: models that would otherwise each need their own signup, contract and payment method, reachable by a string change. It is the fastest way to answer “would a different model do this better” with no procurement at all.

Pros

  • One key and one balance across a wide catalogue, removing per-provider signup entirely
  • Changing model is a string in the model field, so evaluation loops are minutes rather than days
  • A fallback target needing no new vendor relationship, and routing across upstreams for one model can route around a provider incident

Cons

  • A third party between you and the model with prompts crossing it, which becomes a problem the first time you sell to a security-conscious customer
  • Governance is thin: per-team budgets, audit trails and redaction do not live here
  • You inherit whichever upstream served the request, so identical calls can differ in behaviour and latency

Best for: Pre-product-market-fit teams still deciding which model to use, who want the widest access with one account and one invoice.

Pricing: Pay-as-you-go against a prepaid balance, with a margin over the underlying provider’s token rates and no seat fee or commitment.

LiteLLM

LiteLLM homepage

LiteLLM scales with you rather than needing replacement. It is both a library you call in-process and a proxy speaking the OpenAI format, so a startup can start with the library — no extra hop, no deployment — and later run the same translation layer as a service with virtual keys, budgets and spend logs.

Pros

  • Library first, proxy later, with the same provider abstraction either way, so the migration is not a rewrite
  • Broadest provider coverage here, local runtimes included, which matters while you are still experimenting
  • Virtual keys with per-key budgets answer the per-customer cost question without a data project
  • Self-hosted, so prompts never leave your network — the cheapest answer to a prospect’s security questionnaire

Cons

  • The proxy is a stateful service with a database in your request path, which is real work without a platform engineer
  • Provider breadth means uneven edges: well-used providers are solid, rare ones less exercised
  • Large configuration surface, so routing behaviour nobody can fully describe is an easy end state

Best for: Small teams with at least one engineer comfortable running infrastructure, who want a path from in-process library to a real gateway without changing call sites.

Pricing: Open source with no licence cost for the core plus a paid enterprise tier for organisation-level features; self-hosting converts the cost to infrastructure and your own time.

Portkey

Portkey homepage

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks, generally available for enterprises. For a startup that cuts both ways. The product is the same developer-facing gateway — routing, retries, caching, prompt management and tracing behind an OpenAI-compatible endpoint, configured with a config object rather than code. But it now sits inside a large security vendor’s portfolio, which is reassuring if your buyers are enterprises and a roadmap question if you are a small customer of a big company.

Pros

  • The widest feature surface per unit of integration effort: routing, fallback, caching, prompt versioning and tracing in one hop
  • Config-driven routing means changing provider or fallback order needs no deploy
  • Prompt storage decouples iteration from your release process, often the actual trigger for adopting a gateway

Cons

  • Your availability becomes the product of two vendors’, and a startup rarely has the engineering to absorb that
  • As a small account in a large security portfolio, your feature requests compete with a platform strategy
  • Breadth is shallow in places — teams with a serious evaluation practice still run a dedicated tool

Best for: Small teams that want routing, caching, prompt management and tracing without operating anything, and expect enterprise security questions early.

Pricing: Usage-based tiers on requests processed, with a free entry tier for low volume and enterprise contracts above it that are not itemised publicly.

Helicone

Helicone homepage

Helicone solves one problem faster than anything else: nobody can see what the LLM calls are doing. Point the base URL at it and every request, response, token count, latency and error appears in a dashboard, with caching and rate limiting at the same hop. It has announced that it is joining Mintlify. Proxy-first means zero instrumentation, and also that it sees only what crosses it.

Pros

  • The fastest path from zero to per-request logs, with no instrumentation to write
  • Session and user attribution via request headers makes per-customer cost answerable cheaply
  • A self-hostable open-source build removes the objection about prompts leaving your network

Cons

  • Proxy-only visibility: agent steps, retrieval and tool executions never reach it, which matters more the more agentic your product gets
  • Now part of a documentation company’s portfolio, a genuine unknown for a multi-year dependency
  • Routing, fallback and hard budget enforcement are lighter than in policy-built gateways

Best for: Teams whose immediate need is visibility and per-customer attribution rather than routing, and who want it working this afternoon.

Pricing: Usage-based on logged requests with retention tiers and a free entry tier for low volume; the self-hosted build costs only the infrastructure you run it on. Deeper tracing lives in LLM observability tools.

Cloudflare AI Gateway

Cloudflare AI Gateway homepage

The lowest-effort managed hop available: change the base URL to a Cloudflare endpoint and you get logging, caching, rate limiting and retries with nothing deployed and no library added. Because it terminates at the edge, the extra hop is usually short — the most credible answer to the latency objection, and the reason it is a sensible default for a team with no platform engineer.

Pros

  • Nothing to run, nothing to install, no new dependency in your build
  • Edge termination keeps the added hop short, which matters most for time-to-first-token
  • Caching and rate limiting are the first two controls most teams want, and they are built in

Cons

  • A managed third party in the path of every model call, with your prompts crossing it
  • Narrow depth: no serious prompt management, no evaluation, limited policy surface
  • Deepens coupling to one cloud’s ecosystem, a real cost if you are otherwise portable

Best for: Small teams with no platform engineer who want logs, caching and rate limiting in front of provider calls today.

Pricing: Usage-based on requests through the gateway, with a free allowance at low volume and caching and log retention behaviour tied to plan level rather than sold separately.

Vercel AI Gateway

Vercel AI Gateway homepage

If your product is a Next.js application on Vercel using the AI SDK, this is the path of least resistance: one endpoint, many providers, no separate provider accounts, and usage on a bill you already receive. That last point is less trivial than it sounds at seed stage, when every vendor means another card and another invoice.

Pros

  • Removes provider signup and key handling entirely for teams already on the platform
  • Fits the AI SDK’s streaming and tool-calling abstractions, where application code already lives
  • Provider fallback and model switching without a contract with each provider

Cons

  • Strongly coupled to one hosting platform; the value drops the day you move a service elsewhere
  • Governance is not the centre of gravity: hard budgets, audit trails and redaction are thin
  • Another layer of vendor margin between you and the provider as volume grows

Best for: Startups shipping a Next.js product on Vercel with the AI SDK who want model access and fallback with no new vendor relationship.

Pricing: Usage-based on tokens routed through the gateway, billed alongside existing platform usage rather than as a separate contract.

How to choose

Do this in an afternoon, not a sprint.

Check the triggers first. One provider, one service, no per-customer cost question, prompts changed by the person who wrote them? Wrap the SDK, keep the call behind one function, and revisit in three months. Adopting a gateway is cheap; removing one is not.

Pick on the trigger that fired. Cost attribution points to virtual keys per tenant. Visibility points to a proxy that logs by default. Prompt iteration points to prompt storage. A second provider points to routing and fallback — and if that is genuinely the problem, model routing tools go further than any gateway’s default.

Then apply the two constraints that eliminate most candidates. Can prompts leave your infrastructure? If a design partner has already said no, the managed options are gone. Is there anyone to operate a stateful service? If not, a self-hosted proxy is a trap however good it is.

Test the fallback before you rely on it. Force the primary to fail and run your tool-calling and streaming paths against the secondary. A two-hour job that prevents the most common gateway disappointment.

OptionRuns whereTime to a routed requestPicks itself when
OpenAI SDK plus retry wrapper (no gateway)In your processMinutes, no fallbackOne provider, one service, no cost questions yet
OpenRouterManaged aggregatorMinutesYou are still choosing a model
LiteLLMLibrary or self-hosted proxyHours as a library, days as a proxyYou want a path that survives growth, and can run infrastructure
Portkey / Prisma AIRSManagedHoursYou want routing, caching and prompts without operating anything
HeliconeManaged or self-hostedMinutes for logsNobody can see what the calls are doing
Cloudflare AI GatewayManaged edgeMinutesYou have no platform engineer and want the basic controls
Vercel AI GatewayManaged platformMinutesThe product already lives on Vercel

Free tiers exist across the managed options and are useful for evaluation, but read the shape rather than the number: whether the limit is requests, tokens or log retention, whether exceeding it drops data or blocks calls, and whether the features you came to test are even available on the free tier. A gateway whose free tier omits fallback cannot tell you whether its fallback works.

The AI gateways hub covers the architectural decision in more depth, and AI gateways for enterprise is the same shortlist once procurement is involved — worth skimming early if you sell upmarket, because it tells you which of these choices you will have to redo.

Frequently asked questions

Is a base-URL swap really all it takes?

To get a response, yes. To get identical behaviour, no. Provider-specific parameters, tool-call formatting and streaming details are where compatible endpoints diverge, and those differences surface exactly when your fallback fires. Budget an afternoon to verify behaviour rather than connectivity.

Should a two-person team run a self-hosted gateway?

Usually not. It is a stateful always-on service in the path of your most important calls, and at two people nobody is on call for it. Use the library form, or a managed hop, until someone owns infrastructure as a job.

Will a gateway save us money?

Sometimes, through caching and cheaper routing, never automatically. What it reliably buys is knowing where the money goes, which is the precondition for saving any.

How do I keep the option open without adopting one now?

Put every model call behind one function taking a model name, messages and options, and keep provider-specific parameters out of your call sites. Adopting a gateway later is then a one-file change rather than an audit of forty.