Buyer’s Guide

Best GraphQL Servers and Platforms

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • graphql
  • api
  • federation
  • gateway

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

The GraphQL migration that goes wrong almost never goes wrong at the schema. It goes wrong three months later, when someone asks why the p99 on the mobile home screen tripled and nobody can answer, because every request in the access log is a POST to /graphql that returned 200.

That is the shape of the real problem. GraphQL moves query planning out of the database and into your resolver code, moves caching out of HTTP and into your application, and moves rate limiting out of “requests per endpoint” into a cost model you have to invent. Each of those moves is defensible. Making all three at once, in a migration, while also standing up federation across four teams, is how a REST stack that was merely annoying becomes a distributed system that is genuinely hard.

Most teams reading this are mid-migration rather than greenfield. You already have REST endpoints, an existing gateway, clients in production, and a mobile team that wants to stop making six calls to render one screen. The question is not whether GraphQL is good. It is which server you put in front of your existing services, how much of the operational surface it hands back to you, and whether the federation story is worth the control plane it requires.

I have watched this decision from the platform side, and the pattern that repeats is this: teams evaluate GraphQL servers on developer experience and end up buying an operations problem. So this guide leads with the four mechanisms that decide whether your GraphQL layer survives contact with production, spends a full section on the teams who wish they had not done it, and only then names products.

Key takeaways

  • GraphQL’s flexibility is paid for by giving up HTTP’s cheapest features: URL-shaped caching, per-endpoint rate limits, and status codes that mean anything. You rebuild each of those inside your server.
  • The N plus one problem is not a bug you fix once. It is the default behaviour of field resolvers, and a server either batches for you at the data layer or hands you DataLoader and a discipline.
  • Automatic persisted queries and trusted document safelisting are different things that sound identical. One saves bandwidth; only the other stops arbitrary queries from reaching your executor.
  • Federation is a control plane, not a library. If you cannot run composition checks in CI and block a subgraph deploy that breaks the supergraph, you have distributed your schema without distributing the safety.

GraphQL: the query language underneath, not a product

Before comparing servers it is worth separating what GraphQL specifies from what every vendor below adds on top, because a surprising amount of what people call “GraphQL problems” are actually “things the spec deliberately left to you”.

GraphQL is a specification. It defines a type system, a query language, a validation algorithm and an execution algorithm. There is no company selling it, no licence, no bill. It says nothing about HTTP, nothing about caching, nothing about authorization, nothing about how a schema is split across services, and nothing about how you stop a client from asking for a thousand nested nodes.

That silence is the entire reason this category of products exists.

What it gives you

  • A typed contract that is introspectable at runtime, so clients, code generators and documentation tooling read the same source of truth without a side-channel artifact
  • Client-specified response shapes, which is what collapses six round trips on a mobile screen into one and removes the endless “can you add this field to the list endpoint” ticket
  • Validation before execution: a malformed or unknown-field query is rejected by the server’s validation phase without touching a resolver, which is a real safety property you do not get from JSON over REST
  • Deprecation as a first-class schema concept, so a field can be marked dead and still served while you watch who is using it
  • A single execution model across whatever sits behind it, so REST services, databases and gRPC backends compose into one graph

What it does not do

  • It does not define a transport. HTTP semantics, status codes and caching are conventions layered on top, which is why a GraphQL error commonly arrives as a 200 with an errors array and every monitoring tool you own reads it as success
  • It does not define authorization. Field-level access control is your code, in every resolver, forever, unless the server gives you a policy layer
  • It does not bound cost. A valid query can be arbitrarily expensive, and nothing in the spec stops it
  • It does not specify how multiple services form one schema. Federation and stitching are vendor and community inventions with incompatible semantics
  • It does not solve N plus one. Execution is field by field, and the naive implementation is one database call per field per item

Treat the spec as the floor. Everything that makes GraphQL operable in production is above it, and that is what you are actually choosing between.

The four mechanisms that decide whether your GraphQL layer survives

N plus one is the default, not the failure mode

GraphQL executes field by field. Resolve a list of twenty orders, then resolve customer on each one, and the naive implementation issues twenty customer lookups. Nest one level further and it multiplies again. Nothing in the execution algorithm knows that those twenty lookups are one WHERE id IN (...).

There are exactly two strategies, and which one a server uses is the single largest architectural difference in this list.

Batching at the resolver layer. The DataLoader pattern: within a single execution tick, collect every key requested for a given loader, dispatch one batched call, distribute results back. It works, it is well understood, and it is a discipline. Every new resolver is an opportunity to forget it, and the failure is silent until the query shape changes in production. Code review is your only enforcement unless you add query-level assertions in tests.

Compiling the query to the data layer. Servers that own the connection to the database, rather than calling your handwritten resolvers, can translate a GraphQL selection set into one SQL statement with joins and lateral subqueries. N plus one does not arise because there is no N. This is strictly better where it applies and does not apply at all when the data behind a field is a third-party REST call.

Most real graphs are a mix: database-backed regions where compilation works, and service-backed regions where batching is the only option. A server that pretends otherwise is optimising for the demo.

Needs first-hand data: Take your three highest-traffic production query documents and count the actual backend calls each one issues, using database statement logs rather than application traces. Then reshape the same documents into a compiled server and count again. The ratio between naive resolvers, DataLoader batching and query compilation on real query shapes is the number that should decide your server, and nobody publishes it for anything but toy schemas.

Complexity limiting, and why depth limits are not enough

A depth limit is the first thing every team adds and the least useful. It stops the recursive friends { friends { friends } } attack and does nothing about breadth: a flat query requesting forty expensive fields on a page of a thousand items passes a depth limit of five without noticing.

What actually works is static cost analysis. Assign a weight to each field, multiply list fields by their pagination argument, sum the tree before execution, and reject anything over a budget. The important detail is that this happens at validation time, from the query document alone, before a single resolver runs. That is the property that makes it a defence rather than a mitigation.

Two things break naive cost analysis. First, a pagination argument supplied as a variable is not known at static analysis time unless you resolve variables first, so implementations that analyse the raw document undercount. Second, cost is not uniform: a field backed by a cached lookup and a field backed by a cross-region service call cannot share a weight, so the weights need to be assigned per field by someone who knows the backend, and then maintained.

Budgets belong per client identity, not globally. An internal service and an anonymous public client should not draw from the same pool, which is the same argument made in the rate limiting guide and lands the same way: the limit is only meaningful if the identity behind it is.

Persisted queries: two different features with one name

This distinction is worth getting right because it is routinely conflated in vendor material and the security consequence is large.

Automatic persisted queries. The client hashes the query document and sends the hash. If the server does not recognise it, the client resends the full document and the server caches it. The benefit is bandwidth, and the ability to use GET so a CDN can cache the response. The server still accepts and registers arbitrary new documents at runtime, so this is an optimisation, not a control.

Trusted documents, also called safelisting. At build time, the client’s queries are extracted into a manifest, published to the server or gateway, and the server rejects any operation whose hash is not in the manifest. Arbitrary queries cannot reach the executor at all.

Safelisting is the strongest control available for an internal-facing graph, and it changes your cost calculation completely: if the only documents that can execute are ones you shipped, complexity analysis becomes a build-time check rather than a runtime tax, and introspection can be turned off in production without breaking clients.

It also imposes a real constraint. Client and server deploys become coupled, because a mobile app in the wild is still sending last quarter’s document hashes and the manifest must retain them. Any third-party consumer of your graph cannot use it at all. Which is the honest reason safelisting is common on internal graphs and rare on public ones.

Federation is a control plane

Federation splits one schema across independently deployed services. Each subgraph owns some types and contributes fields to shared entities; a composition step merges them into a supergraph schema; a router accepts client queries, builds a query plan that fans out to subgraphs, and stitches the results.

The library part of that is easy. The part that decides whether it works is the control plane:

  • Composition checks in CI. A subgraph change that breaks composition, or removes a field another subgraph depends on, must fail the pull request. If composition only happens at deploy time, you have shipped a distributed monolith where any team can break the graph.
  • A schema registry with usage data. Before you remove a deprecated field you need to know which client operations still select it. Without per-field usage attribution, deprecation is permanent and the schema only grows.
  • Query plan visibility. When a federated query is slow, the useful question is which subgraph fetch in the plan was slow and whether the plan required a sequential hop that could have been parallel. If the router does not expose the plan, you are debugging blind.

Needs first-hand data: Instrument your router to emit the query plan alongside per-subgraph fetch timings for your top ten federated operations, then report what fraction of total wall-clock time is sequential entity resolution rather than parallel fetches. That fraction is the real cost of your entity boundaries, and it is the number that tells you whether a type is split across the wrong services.

Entity resolution is where the sharp edge lives. A shared entity is fetched from one subgraph by key and enriched by another, which means a single client query becomes a sequential dependency chain across services, and the N plus one problem reappears at the network layer where DataLoader cannot help you. Routers batch entity fetches for exactly this reason, and how well they do it is a genuine differentiator.

The teams who regretted GraphQL

This is the section most GraphQL comparisons omit, and it is the one worth reading. I have seen migrations reversed. The reasons are consistent, and every one of them is predictable in advance.

You gave up CDN caching and did not notice until the bill arrived. REST responses cache at the edge on a URL. GraphQL’s default is a POST to a single path, which no CDN will cache. You can recover this with persisted queries sent over GET, and most teams do not do it, and the result is that every request that used to be served from an edge node now reaches your origin. If a meaningful share of your traffic was public and cacheable, GraphQL turns a cheap CDN problem into an expensive origin problem.

Your observability went blind on day one. Everything is one route, one method, and status 200 whether it worked or not. Endpoint latency dashboards, per-endpoint error rate alerts, gateway rate limits keyed on path, WAF rules keyed on URL: all of them stop distinguishing anything. You get this back only by instrumenting at the operation level, exporting operation name and field-level timings as span attributes, and rebuilding your alerts around them. That is real work, it happens after the migration, and until it lands your on-call has less visibility than they had with REST. This is the same problem covered from the monitoring side in the API monitoring guide, and GraphQL is its hardest case.

Authorization got distributed across hundreds of resolvers. In REST, an endpoint is a natural authorization boundary. In GraphQL, any field can be reached from any query path, so “can this user see this field” has to be answered at the field, in every path that reaches it. Teams that did not build a policy layer early ended up with authorization logic scattered through the resolver tree, and no way to answer “who can read salary” other than reading code.

One team, one client, one backend. GraphQL’s payoff is proportional to the number of distinct clients with different data needs and the number of backends being composed. If you have a single web frontend talking to a single service, you have taken on query planning, cost analysis, caching and a new tier for essentially no gain. A well-designed REST endpoint that returns exactly what the one screen needs is the correct answer, and it is not a less sophisticated answer.

The schema became append-only. GraphQL replaces versioning with deprecation, which is genuinely better, but only if you can eventually delete. Deleting requires knowing that no live client selects the field, which requires per-field usage tracking tied to client identity. Teams without a registry that collects that data end up unable to remove anything, and five years later the schema carries every field anyone ever added.

Needs first-hand data: Before migrating a public surface, measure what share of your current REST traffic is served from the CDN edge rather than reaching origin, broken down by endpoint. Then model the same traffic as GraphQL POSTs with no edge cache. The difference in origin request volume is the single largest hidden cost of a GraphQL migration and it is knowable from your existing CDN logs today.

The escape hatch that works, when it comes to that: keep GraphQL as an internal backend-for-frontend that your own clients consume with safelisted documents, and expose REST or gRPC to anyone outside that boundary. You keep the composition benefit where it pays and you stop exporting your executor to the internet. That split is also the cleanest way to keep a conventional gateway in front, which is where the API management platform layer still applies.

Apollo GraphQL

Apollo GraphQL homepage

Apollo is the centre of gravity in this category and the reference implementation of federation. The current architecture separates concerns cleanly: subgraphs can be any spec-compliant server, Apollo Router is a Rust query planner and executor that sits in front, and GraphOS is the hosted control plane holding the schema registry, composition checks, usage reporting and persisted query manifests. If you adopt federation, you are adopting Apollo’s semantics whether or not you buy Apollo’s platform, because the federation directives are theirs.

Pros

  • The most complete federation control plane available: composition checks that run in CI, schema change proposals, and per-field usage attribution tied to client identity, which is what makes deprecation finishable
  • Apollo Router’s query planner is the most battle-tested here, with entity fetch batching and exposed query plans that make slow federated queries debuggable rather than mysterious
  • Trusted document safelisting is a first-class workflow with a manifest published from client builds, not something you assemble yourself
  • The client side of the ecosystem is mature enough that normalized caching, code generation and fragment colocation are solved problems rather than projects

Cons

  • Federation’s directive set is Apollo-defined, so a schema built around it is portable only to routers that implement Apollo’s semantics, and that is the lock-in that matters
  • The genuinely valuable parts, meaning the registry, usage reporting and schema checks, live in the hosted platform, so the open-source path gives you the runtime without the control plane that justifies federation
  • Operating a router plus subgraphs plus a registry is a multi-component platform commitment, and small graphs pay that overhead for benefits they will not use
  • Router customisation pushes you toward Rust plugins or a coprocessor hop, which is a steeper extension story than a Node middleware chain

Best for: Organisations with several teams owning separate subgraphs who need composition checks and field-level usage data before they can safely let the schema evolve.

Pricing: Open-source router and server components with no licence cost, plus a hosted platform tier metered on operations reported to the registry with enterprise agreements for larger deployments and self-managed control planes.

Hasura

Hasura homepage

Hasura inverts the usual model. Instead of you writing resolvers, it introspects your database and generates a GraphQL API, then compiles incoming queries into a single SQL statement rather than executing field by field. The N plus one problem does not occur in the database-backed region of the graph because the whole selection set becomes one query. Authorization lives in declarative metadata as row and column permissions evaluated against session variables, which keeps it out of resolver code.

Pros

  • Query compilation to a single SQL statement removes the entire class of N plus one failures for database-backed fields, and does it without any discipline on your part
  • The permission model is declarative metadata rather than code, so “who can read this column” is answerable by reading configuration instead of auditing resolvers
  • Remote schemas and actions let handwritten services and third-party APIs join the same graph, so the generated portion does not have to be the whole graph
  • Subscriptions are implemented as multiplexed polling of the compiled query, which scales far better than a naive per-connection subscription implementation

Cons

  • A schema generated from tables is a database schema, not a product API, and shipping it directly exposes your storage model to clients in a way that is very hard to walk back
  • The permission engine is powerful but is its own language, and complex authorization rules become metadata that is harder to test than the code it replaced
  • The major version transition reshaped the engine and the metadata model, so material written for the earlier generation does not map cleanly and migration planning is a real project
  • Open-source and commercial boundaries have moved over the product’s life, which is a genuine continuity risk to weigh if self-hosting is your compliance answer

Best for: Teams whose graph is mostly Postgres-backed and who want correct data-layer performance and declarative row-level authorization without writing a resolver tier.

Pricing: Open-source engine with no licence cost when self-hosted, plus a hosted tier metered on data passthrough and active models, with enterprise agreements adding governance and support.

WunderGraph

WunderGraph homepage

WunderGraph’s Cosmo is the credible open-source alternative to a commercial federation control plane. It implements a federation-compatible router, a schema registry, composition checks and analytics, and the whole stack can be self-hosted under a permissive licence. For teams that want federation’s benefits but cannot accept the registry living in someone else’s cloud, this is the shortest path.

Pros

  • A full federation control plane, registry and composition checks included, that you can run entirely inside your own infrastructure with no licence cost
  • Federation-compatible semantics mean existing subgraphs generally move across without reshaping the schema, which makes it a real migration target rather than a rewrite
  • Router analytics and query plan visibility are included rather than gated, so the debugging surface is available on the self-hosted path
  • The broader WunderGraph lineage is built around persisted operations, so safelisting is treated as the normal way to run a graph rather than an advanced option

Cons

  • Smaller ecosystem and community than Apollo, which means fewer people have already hit your problem and fewer integrations exist off the shelf
  • You are tracking compatibility with a specification controlled by a competitor, so a semantic change upstream is a roadmap risk you do not control
  • Self-hosting the control plane means operating a registry, a router fleet and an analytics store, which is the platform work the hosted option exists to avoid
  • Product direction has shifted across the company’s life from a BFF framework to a federation platform, so older documentation and examples describe a different product

Best for: Platform teams that want federation with a registry and composition checks running inside their own network rather than a vendor’s.

Pricing: Open source under a permissive licence with no cost to self-host, plus a managed cloud tier metered on router traffic and registry usage.

GraphQL Yoga

Yoga is a spec-compliant GraphQL server from The Guild, built on a plugin architecture that makes the request lifecycle explicit: parsing, validation, execution and result formatting are all hookable. It runs on Node, Deno, Bun and Workers runtimes from the same code, which matters more than it sounds if part of your graph belongs at the edge. It is a server, not a platform, and that is the point.

Pros

  • The plugin system makes cost analysis, safelisting, response caching and tracing composable additions rather than forks of the server
  • Genuine portability across JavaScript runtimes, including edge workers, so the same server code deploys where your traffic is
  • No vendor coupling anywhere in the stack: the schema is plain GraphQL, and the federation and registry pieces are separate choices you can make independently
  • Small enough to reason about completely, which is worth a lot when you are debugging execution behaviour at two in the morning

Cons

  • Nothing is included. N plus one protection, cost limits, safelisting and usage reporting are each a decision and an integration, and a team that does not make them ships an unprotected graph
  • No control plane at all, so federation, registry and schema checks require assembling separate components and owning the seams
  • Being a library rather than a product means there is no escalation path when something is wrong in production beyond your own engineers and a community issue tracker
  • Resolver-level performance is entirely on you, so the data-layer efficiency that a compiling server gives for free has to be engineered

Best for: Teams building a single deliberate GraphQL service who want an unopinionated, portable server and intend to choose each production control themselves.

Pricing: Open source with no vendor and no bill. Your only cost is the infrastructure you run it on and the engineering time to assemble the controls it deliberately leaves out.

Grafbase

Grafbase is a federation gateway and platform positioned around edge deployment: the router runs close to clients, with a schema registry, composition checks, caching and analytics attached. The pitch is that a federated graph is a latency-sensitive component and belongs at the edge rather than in one region, which is a reasonable argument if your subgraphs are themselves distributed.

Pros

  • Edge-first router placement removes a round trip for clients far from your primary region, which is the one latency win a gateway tier can actually deliver
  • Registry, composition checks and analytics ship together, so the control plane is not a second purchase
  • Response caching is treated as a built-in concern rather than something you bolt on, which addresses GraphQL’s worst structural weakness directly
  • Federation compatibility means existing subgraphs are usable without reshaping the schema

Cons

  • Putting the router at the edge only helps if subgraphs are reachable quickly from the edge; with a single-region backend you have added a hop and moved the latency, not removed it
  • A younger platform than the incumbents, so ecosystem depth, integrations and third-party operational knowledge are thinner
  • The managed control plane is the product, which makes a fully self-hosted deployment a weaker story than the open-source alternatives here
  • Caching a personalised graph safely requires per-field cache scoping that is easy to get subtly wrong, and the failure mode is serving one user another user’s data

Best for: Teams with a globally distributed client base and distributed subgraphs who want the federation router and its control plane as one managed edge service.

Pricing: Usage-based metering on gateway requests with a free entry tier, and enterprise agreements for larger or self-managed deployments.

PostGraphile

PostGraphile reflects a Postgres schema into a GraphQL API and, like Hasura, compiles queries into efficient SQL rather than executing resolvers per field. Its distinguishing choice is that authorization is Postgres row-level security: the server sets session variables and the database enforces the policy. That means your access rules live in one place, enforced by the engine that owns the data, and they apply identically to anything else that connects.

Pros

  • Delegating authorization to Postgres row-level security means the policy is enforced at the data layer, so a bug in the API tier cannot bypass it
  • Query compilation produces genuinely efficient SQL including lateral joins, so nested selections do not degenerate into per-row queries
  • A deep plugin system lets you rename, hide, wrap and extend the generated schema, which is what makes the generated API survivable as a product surface
  • No hosted dependency at all: it is a library in your own Node process with no control plane to buy or operate

Cons

  • Postgres only, which rules it out the moment a meaningful part of your graph is a third-party service or a non-Postgres store
  • Row-level security is a powerful and unforgiving tool, and policies that are subtly wrong fail closed in ways that are hard to debug and fail open in ways that are worse
  • The generated schema still reflects your tables, so exposing it publicly couples external clients to internal storage decisions unless you invest in the plugin layer
  • Federation, registry, usage tracking and safelisting are all outside its scope, so a multi-service graph needs another product on top

Best for: Postgres-centric teams who want a high-performance generated GraphQL API with authorization enforced by the database rather than the application tier.

Pricing: Open source with no licence cost for the core, with commercial plugin and support offerings available from the maintainers for production features.

StepZen

StepZen takes the declarative route: rather than writing resolvers, you describe how GraphQL types map onto existing REST endpoints, databases and other backends using directives, and the engine executes the plan. It is aimed squarely at the case this article opened with, a team with existing REST services that wants a GraphQL surface without building a new service tier. StepZen is now part of IBM’s API portfolio, which is the most important fact about it.

Pros

  • Declarative backend mapping means a REST-backed graph can be assembled configuration-first, without a resolver codebase to maintain
  • Built-in request batching and sequencing across declared backends addresses N plus one for REST-backed fields, which is the case compiling database servers do not cover
  • Composing multiple existing APIs into one graph is the design centre rather than an add-on, which fits a migration better than a greenfield server does
  • Enterprise integration with a larger API management portfolio gives procurement a single relationship rather than another vendor

Cons

  • Acquisition into a large enterprise portfolio makes independent product trajectory an open question, and that is a serious consideration for a component this central
  • Declarative configuration is elegant until a backend needs logic that the directive set does not express, and the escape hatch is less pleasant than ordinary code
  • The federation and registry story is weaker than the dedicated platforms here, so a large multi-team graph is not its strength
  • Community and third-party material are thin compared to the open-source options, which raises the cost of every unusual problem

Best for: Enterprises with a large estate of existing REST services who want a declarative GraphQL layer over them inside an established API management relationship.

Pricing: Sold as part of an enterprise API management portfolio with tiering and agreements negotiated per deployment rather than a public self-serve meter.

How to choose

Answer three questions in order and most of the list disappears.

Is your graph mostly one database, or mostly many services? Mostly one database, and a compiling server (Hasura or PostGraphile) gives you correct data-layer performance for free and removes the largest failure mode in the category. Mostly services, and compilation does not apply, so you are choosing a server plus a batching discipline plus a composition strategy.

Do you need more than one team to own parts of the schema? If no, do not adopt federation. A single subgraph with a router in front is pure overhead. If yes, you need a control plane with composition checks in CI, and your real choice is Apollo’s hosted platform, Cosmo self-hosted, or Grafbase managed at the edge.

Can your graph be safelisted? If every client is yours, safelist trusted documents and turn introspection off in production. That single decision removes arbitrary query cost, most of the complexity-limiting problem and a large share of the attack surface. If you have third-party consumers, you cannot, and cost analysis with per-client budgets becomes mandatory rather than optional.

Then, before you commit, instrument one real screen end to end. Count backend calls, not response times. The gap between what you think your resolvers do and what they actually do is the whole decision.

OptionShapeN plus one strategyControl planePicks itself when
GraphQL(standard, not a product)UnspecifiedNoneNever; it is the layer everything else implements
Apollo GraphQLRouter plus hosted platformDataLoader in subgraphs, entity batching in routerFull, hostedSeveral teams own subgraphs and deprecation must finish
HasuraGenerated, compiling engineQuery compiled to one SQL statementMetadata and consoleThe graph is mostly Postgres and authorization is declarative
WunderGraphFederation router plus registryEntity batching in routerFull, self-hostableFederation is required but the registry cannot leave your network
GraphQL YogaLibrary serverYour DataLoader disciplineNoneOne deliberate service, every control chosen by you
GrafbaseManaged edge gatewayEntity batching plus response cachingFull, managedClients are global and caching is the binding constraint
PostGraphileGenerated, compiling libraryQuery compiled to one SQL statementNonePostgres only, with row-level security as the policy engine
StepZenDeclarative composition engineDeclared batching across REST backendsEnterprise portfolioA large REST estate needs a configuration-first graph

Frequently asked questions

Do I need federation, or is schema stitching enough?

Federation exists because stitching put the merge logic in the gateway, which made the gateway a shared codebase every team had to change. Federation moves ownership into subgraphs and leaves the router mechanical. If you have one team, neither is worth it. If you have several, federation’s value is not the merge, it is that composition can be checked in CI and a bad change fails in a pull request instead of at deploy.

How do I rate limit GraphQL when every request hits the same endpoint?

Not by request count. Compute a static cost for the query document during validation, then charge that cost against a per-client budget, the way a token bucket works but with query cost as the unit. Request-count limits on /graphql are worse than useless because they treat a trivial query and a thousand-node query identically. If your documents are safelisted you can precompute the cost of every known operation at build time and skip the runtime analysis entirely.

Can I put a normal API gateway in front of a GraphQL server?

Yes, and you should for TLS, authentication, IP controls and basic protections, but understand what it cannot do. A conventional gateway sees one path and one method, so path-based routing, per-endpoint quotas and URL-shaped WAF rules do nothing. Anything that needs to distinguish one operation from another has to live in the GraphQL server or a GraphQL-aware router.

Is GraphQL a bad choice for a public API?

It is a demanding one. Public means untrusted documents, which means you cannot safelist, which means cost analysis, introspection policy, per-client budgets and abuse monitoring are all mandatory and permanent. Several large public GraphQL APIs run well with exactly that investment. The question is whether you want that to be a permanent line item, or whether REST for outside consumers and GraphQL for your own clients gets you most of the benefit for a fraction of the operational surface.

What breaks first when a GraphQL migration goes badly?

Caching economics and on-call visibility, in that order, and both show up weeks after launch rather than at cutover. Plan the operation-level instrumentation and the persisted-query-over-GET path as part of the migration, not as follow-up work, because the follow-up work is what gets cut when the migration runs long.