Shipping an SDK is not a code generation problem. It is a distribution commitment, and teams consistently underestimate which half is the work.
The generation part takes an afternoon. What follows is permanent: six package registries with six release processes, semantic versioning decisions on every spec change, changelogs, deprecation notices, a support surface in six languages, and the fact that every consumer now has a pinned version of your API surface embedded in their build. An SDK is the most durable artifact you publish, because unlike documentation, people compile against it.
Which reframes the evaluation. The interesting question is not “does it produce a TypeScript client” because everything does. It is: when you rename one field in the spec, what lands in the pull request, what version number does it get, and does the consumer’s build break loudly or silently accept a type that is now wrong?
The second question is about quality, and quality in this category means something specific. A generated client is good when a developer who has never seen your API can guess the method name, gets useful types in their editor, receives a typed error instead of a parsed string, and never has to read your reference documentation to make a paginated request. That standard is met by a minority of generated SDKs, and the gap between the tools here is larger than in any other category in this cluster.
Key takeaways
- Generated code quality varies enormously by target language, and a tool that is excellent in TypeScript can be mediocre in Go. Evaluate per language you will actually ship.
- The publish pipeline is the product. Generation is an afternoon; versioning, releasing to six registries and maintaining changelogs is forever.
- Output quality tracks spec quality. Operation identifiers, named schemas and correct nullability determine whether a generated client is usable in any generator.
- A spec change that is not breaking for the API can still be breaking for the SDK, because method names, type names and required arguments come from the document.
What “good” means, and why it is different per language
The generic version of this is useless, so here is the concrete version. A generated client is judged on seven things, and the weight of each differs by ecosystem.
Method naming. Does the client expose client.orders.list() or getOrdersOrdersGet()? This comes almost entirely from operation identifiers in the spec and from whether the generator groups operations into resource namespaces. It is the first thing a developer sees and the thing that decides whether the SDK feels designed or extruded.
Type fidelity. In TypeScript, Go, Java, Kotlin, C# and Rust, the types are the main reason to use an SDK at all. The specific question is nullability: if the spec says a field is optional, does the generated type make it optional, and does that force the caller to handle absence? Generators that mark everything optional defensively produce types that are technically accurate and practically useless, because every field access becomes a null check the developer stops thinking about. In Python and Ruby the stakes are lower but not zero, since type hints drive editor completion.
Error handling. A good client raises a typed error carrying the status code and the parsed error body. A bad one raises a generic HTTP exception with a string, so every consumer writes their own parsing, differently. This depends entirely on whether your spec documents error response schemas, which most do not.
Pagination. Every API has pagination and every consumer writes the same loop. A good SDK exposes an iterator or an async iterator that handles it, so the consumer writes a for loop and never sees a cursor. This is the single highest-value ergonomic feature and a clear dividing line between the purpose-built generators and the template-driven ones.
Retries, timeouts and idempotency. Sensible defaults with per-call overrides, retrying only on the statuses and methods where retrying is safe. Generated clients that retry a non-idempotent POST on a 500 will eventually create duplicate charges for somebody.
Streaming and long-running operations. Server-sent events, chunked responses, file uploads and polling-based jobs. These are where generated clients most often fall back to “here is the raw response, good luck”, which is fine as an escape hatch and bad as the only option.
Idiomatic feel. Does it look like code a person from that community would write? Context as the first Go argument. Async-first in Python with a sync wrapper. Builders where the language expects builders. This is the hardest thing to automate and the clearest separator between generators covering six languages deeply and generators covering fifty shallowly.
That last point is the structural tradeoff in this whole category. Template-driven generators achieve breadth by treating languages as mostly interchangeable with different syntax. Purpose-built generators achieve quality by having someone who writes that language for a living own each target. You cannot have both, and which you need depends on whether your SDKs are a developer-facing product or an internal convenience.
Needs first-hand data: Take one real spec, generate clients with two or three candidates for the same two languages, and hand them to an engineer who has not seen your API with one task: list a paginated resource, handle a 429, and parse a validation error. Record time to working code and questions asked. That test discriminates between generators far better than reading the output does, because output that looks fine can still be unguessable.
The publish pipeline is the actual product
Generation is the demo. This is the job.
Versioning decisions. Every spec change needs a version number, and the mapping from API change to SDK version is not one-to-one. Adding an optional response field is non-breaking for the API and non-breaking for most SDKs. Renaming an operation identifier is completely non-breaking for the API and a compile error in every consumer of the SDK, because it renames a method. That asymmetry is the thing to internalise: SDK breakage and API breakage are different sets. Anything automating this needs to classify changes against the generated surface, not the HTTP surface.
Registry mechanics. npm, PyPI, Go modules, Maven Central, NuGet, RubyGems, crates.io. Each has its own authentication, its own publishing quirks and its own conventions about namespaces and deprecation. Go modules in particular require a major version suffix in the import path, which makes a major bump a repository-structure change rather than a version string. Maven Central has a signing and staging process that surprises everyone the first time. This work is not hard, it is just relentless, and it recurs on every release.
Changelogs and migration notes. Generated diffs produce accurate and unreadable changelogs. Somebody has to write the sentence explaining what a consumer needs to do. Tools that generate a per-language migration note from a classified diff save real time here; tools that dump a list of type changes do not.
Handwritten code alongside generated code. Every SDK eventually needs something the generator cannot produce: a convenience helper, a webhook signature verifier, a file upload wrapper, a retry policy specific to one endpoint. The mechanism the tool offers for this determines whether your SDK stays maintainable. Per-file overrides and designated extension points survive regeneration. Patches applied on top of generated output conflict on every regeneration and eventually somebody stops regenerating, which is how an SDK dies.
Who owns the pull request. The best-shaped pipeline is: spec changes on merge, the generator opens a pull request per language repository with the regenerated code and a proposed version bump, CI builds and tests each one, and a human approves. Fully automatic publishing without a human step will eventually push a broken client to a registry, and registries are unforgiving about republishing. Where that pipeline runs is a CI/CD platform decision.
Needs first-hand data: Instrument one release cycle end to end and record human minutes spent per language on: reviewing the regenerated diff, deciding the version, writing the changelog entry, and fixing the publish when it fails. Multiply by your release frequency and language count. That annual number is the honest budget for shipping SDKs, and it is what tells you whether a commercial generator is expensive or cheap.
What a spec change does to the output
This is the test worth running before committing to anything, because the answers differ sharply and none of the marketing pages discuss it.
Renaming a response field. Should rename one property on one type and nothing else. Poor generators rename the type too if the schema was inline and anonymous, producing a diff across dozens of files for a one-field change.
Renaming an operation identifier. Renames a method. Breaking for every consumer, non-breaking for the API. A tool that flags this as a major SDK change is doing the job; one that ships it as a patch is actively dangerous.
Adding an optional request parameter. Should add an optional argument in a position that does not break existing calls. In languages without named or optional arguments this is where generators get creative, and where the creativity sometimes breaks source compatibility for no reason.
Adding an enum value. The subtle one. If the generated enum is closed, an unknown value received from the server is a deserialisation failure, so adding a value to your API breaks every deployed old client at runtime. Well-designed generators produce open enums with an unknown case. This single decision has caused more production incidents than any other item on this list, and it is worth checking explicitly for every language you ship.
Reordering anything in the spec. Should produce no diff. Some generators emit output ordered by document position, so moving a schema definition produces a large meaningless diff that hides the real change and trains reviewers to skim.
Regenerating with no spec change at all. Should produce an empty diff. Generators that embed timestamps, versions or non-deterministic ordering fail this, which makes every regeneration noisy and makes “is this change intentional” unanswerable.
Determinism and diff minimality are the properties to test first, because they are cheap to check and they determine whether anyone will ever review a generated pull request properly. Run the generator twice on an unchanged spec and diff the output. If it is not empty, that tool will make every future review worthless.
The upstream implication is that your spec needs the same treatment as source code: linting for consistent operation identifiers, named and reused schemas, correct nullability, and a breaking-change gate before it reaches the generator. That is the subject of OpenAPI and Swagger tooling, and it is a prerequisite rather than a nice-to-have. Every generator below produces better output from a better document.
Stainless

Stainless is a commercial, purpose-built SDK platform that generates clients for a focused set of languages and optimises hard for output that reads as handwritten. The model is a configuration layer on top of your OpenAPI document: you describe resource grouping, method naming, pagination shape and other ergonomics in a config, and the generator applies that consistently across every language. It is the tool behind several widely-used developer platform SDKs, which is a meaningful signal in a category where most output is visibly generated.
Pros
- Output quality is the highest end of this category, with resource-grouped methods, real pagination iterators and typed errors rather than generic exceptions
- A configuration layer over the spec lets you fix ergonomics without polluting the OpenAPI document with generator-specific extensions
- Handles the publish pipeline across registries, including version decisions and release pull requests, which is the recurring cost
- Consistent shape across languages, so documentation examples and support answers translate between them
Cons
- Commercial and positioned for companies whose API is the product, so it is hard to justify for internal SDKs
- Focused language set by design, so an unusual target is simply not available
- Deep dependency on a vendor for an artifact your consumers compile against, which is a continuity question worth asking explicitly
Best for: Companies whose public API is a primary product surface, where SDK quality is part of the buying experience and the language set is mainstream.
Pricing: Commercial subscription scaled by languages generated and API surface, sold as a platform rather than per seat.
Speakeasy

Speakeasy occupies similar ground to Stainless, generating idiomatic SDKs across a focused language set from an OpenAPI document, with a strong emphasis on the surrounding automation: release pipelines, version management, changelogs, documentation snippets and generated Terraform providers alongside the clients. That last item is unusual and genuinely valuable for infrastructure products, where a Terraform provider is a maintenance burden of the same shape as an SDK and usually hand-written.
Pros
- Publish automation is a first-class part of the product, covering registry releases, versioning and changelogs rather than only generation
- Terraform provider generation from the same spec removes a separate hand-maintained artifact for infrastructure products
- Generated documentation snippets come from the same pipeline as the SDKs, so code examples in your docs cannot disagree with the client
- Spec linting and suggestions surface the document problems that would otherwise produce poor output
Cons
- Commercial, with cost scaling as you add languages, which is exactly when SDK maintenance was starting to hurt
- Focused language coverage, so unusual targets fall back to a different tool and a different quality level
- The value concentrates in the automation, so a team that only needs one language and releases rarely is overpaying
Best for: Developer-tools and infrastructure companies shipping SDKs in several languages, plus a Terraform provider, on a regular release cadence.
Pricing: Commercial subscription scaled by number of generated targets and API surface. The documentation snippet side connects to API documentation tools.
Fern

Fern generates SDKs and documentation from either an OpenAPI document or its own more opinionated API definition format. That second input is the distinguishing choice: its own format is stricter about the things that make generation good, notably error types, pagination shape and discriminated unions, so it can produce better clients than a permissive OpenAPI document supports. The generators are open source, with a commercial platform around publishing and docs.
Pros
- A stricter definition format captures the things OpenAPI expresses weakly, especially error types and pagination, producing better clients as a result
- Open-source generators mean the code-producing part can be run in your own pipeline without a vendor dependency
- Documentation and SDKs come from the same definition, so snippets and clients stay aligned by construction
- Can consume plain OpenAPI when you are not ready to adopt its own format, so adoption is incremental
Cons
- Its own definition format is another artifact to maintain unless you commit to it as the source of truth, and committing means a conversion project
- Smaller ecosystem than the template-driven incumbents, so unusual requirements have fewer worked examples
- The split between open-source generators and the commercial platform means the free path covers less of the publish pipeline than the demo suggests
Best for: Teams willing to adopt a stricter API definition in exchange for better generated clients and aligned documentation, particularly if they want to keep generation in-house.
Pricing: Open-source generators with no licence cost, plus commercial tiers for the hosted documentation and publishing platform, scaled by API and seats.
openapi-generator
openapi-generator is the breadth option and the default when a commercial tool cannot be justified. Template-driven, covering dozens of client languages and several server frameworks, with a template override mechanism for fixing what you dislike. For internal SDKs, for unusual target languages, and for teams who need something working today with no procurement, it remains the correct answer despite its well-known rough edges.
Pros
- Language coverage far beyond any commercial tool, including targets no vendor will ever support
- Template overrides let you improve output incrementally without forking, which is how teams make it acceptable
- No vendor relationship, so the whole pipeline runs inside your own build with no external dependency
- Also generates server stubs, useful for scaffolding services from an agreed design
Cons
- Output quality varies sharply between targets, and the weaker ones produce clients that are recognisably machine-made and unpleasant to use
- Generator behaviour changes between releases, so pinning is mandatory and upgrades produce large unrelated diffs
- Ergonomics that make purpose-built SDKs good, pagination iterators, typed errors, open enums, are inconsistent or absent depending on target
- JVM dependency in the pipeline, which is friction for teams with no other Java toolchain
Best for: Internal SDKs, unusual target languages, and teams who need clients now without procurement.
Pricing: Open source with no vendor and no bill. The real cost is engineer time on templates, version pinning and the publish pipeline it does not provide.
Kiota
Kiota is Microsoft’s client generator, built around a fluent request-builder API rather than flat method names: you navigate the URL structure in code, so client.users["id"].messages.get() mirrors the path. That model handles very large API surfaces well, because there is no enormous flat client class, and it is the reason it exists, having been built for a genuinely large API. It generates for the languages Microsoft cares about, with consistent shape across them.
Pros
- Fluent path-based builder scales to very large API surfaces without producing an unmanageable single client type
- Consistent design across supported languages, so the mental model transfers between them
- Can generate a client for only a selected subset of endpoints, which keeps client size down for consumers who use a fraction of a large API
- Backed by an organisation with a permanent operational reason to keep it working
Cons
- The request-builder style is unfamiliar and some developers find it less discoverable than named resource methods
- Language coverage follows Microsoft ecosystem priorities rather than the wider market
- Opinionated abstractions mean the generated client brings its own core libraries, which is a dependency your consumers inherit
Best for: Very large API surfaces, particularly in .NET-centric organisations, where selective endpoint generation matters.
Pricing: Open source with no vendor charge. The real cost is the abstraction layer your consumers take on and the narrower language set.
oazapfts
oazapfts is the minimal option: a TypeScript-only generator that turns an OpenAPI document into a small, typed client with no runtime framework and no ceremony. It does one language, does not try to own your publish pipeline, and produces output small enough to read. When the requirement is a typed internal client for a TypeScript service calling another TypeScript service, the heavier tools are solving problems you do not have.
Pros
- Output is small, readable and reviewable, so a generated diff is something a person can actually check
- No runtime framework or client abstraction imposed, so the generated code drops into an existing codebase without dependencies
- TypeScript-only focus means the types are the priority rather than an average across languages
- Trivial to run in a build step, which makes regeneration on every spec change a non-event
Cons
- One language, so a multi-language SDK story needs a second tool and accepts inconsistent shape between them
- No publish automation, versioning or changelog support, so anything customer-facing means building that yourself
- Small project with limited ergonomic features, so pagination and retries are your problem
Best for: Internal TypeScript-to-TypeScript clients where a typed fetch wrapper is the whole requirement.
Pricing: Open source with no vendor and no bill. The real cost is that it covers generation only, so a public SDK still needs a pipeline you build.
How to choose
One: decide whether the SDK is a product or a convenience. Public SDKs that developers evaluate you by justify commercial tooling, because quality is part of the buying experience. Internal SDKs almost never do, and spending on them is the most common overinvestment in this category.
Two: list the languages you will actually ship and support. Not could. Will. Every language is a support surface, a release process and a set of idioms someone has to own. Three good SDKs beat seven neglected ones, and the language list eliminates most tools immediately.
Three: run the determinism test. Generate twice on an unchanged spec and diff. Then rename one field and inspect the diff. Then add an enum value and check whether the enum is open or closed in each target language. These three checks take an hour and discriminate better than anything else available to you.
Four: work out who owns the publish pipeline. If nobody has time to own six release processes, buy the automation. If one language ships once a quarter, do not.
Five: fix the spec first regardless. Operation identifiers, named schemas, documented error responses, correct nullability. Every tool here produces better output from a better document, and this work is not wasted on a later tool change.
| Tool | Languages | Publish automation | Source of truth | Picks itself when |
|---|---|---|---|---|
| Stainless | Focused, mainstream | Included | OpenAPI plus a config layer | The public API is the product and quality is the differentiator |
| Speakeasy | Focused, mainstream | Included, plus Terraform | OpenAPI | You ship several SDKs and an infrastructure provider on a cadence |
| Fern | Focused | In the commercial tier | Its own definition or OpenAPI | You will adopt a stricter definition for better output |
| openapi-generator | Very broad | None | OpenAPI | You need an unusual language, or no procurement is possible |
| Kiota | Microsoft-centric set | None | OpenAPI | The API surface is very large and the org is .NET-centric |
| oazapfts | TypeScript only | None | OpenAPI | An internal typed client is the entire requirement |
The pattern that fails is the middle: adopting a broad template generator for a public SDK, shipping seven languages, and neglecting six of them. Pick two or three languages and make them genuinely good, then add more when consumers ask, and let everyone else use the OpenAPI document directly with whatever generator suits them.
Where SDK generation sits relative to spec, documentation, testing and gateway decisions is mapped in API management platforms.
Frequently asked questions
Should I publish SDKs at all, or just an OpenAPI document?
Publishing the document costs nothing and lets any consumer generate a client in any language, which is a perfectly respectable answer for internal and partner APIs. Publish SDKs when the API is a product and time-to-first-successful-call is part of how you win, or when your API has real ergonomic complexity: pagination, retries, webhook signature verification, long-running jobs. If your API is twelve straightforward endpoints, a good document and a curl example may genuinely serve developers better than a package nobody maintains.
How do I version an SDK relative to the API?
Keep them separate and say so clearly. The SDK follows semantic versioning against its own surface, because that is what consumers compile against, and the API version is a field in the client configuration or the base URL. Tying the SDK version to the API version forces major bumps for changes that break nothing, and worse, hides SDK-breaking changes like a renamed method inside a version number that looks safe.
What happens to consumers pinned to old SDK versions?
They keep working against the API as long as you honour your deprecation policy, which is an API-side commitment rather than an SDK one. The SDK question is how long you will patch old major versions for security issues, and the honest answer for most teams is one previous major and no further. State it in the readme rather than leaving it implicit. API versioning and deprecation covers the API side.
Can I hand-edit generated code?
Only through an extension mechanism the generator understands. Editing generated files directly works exactly once, and then the next regeneration either overwrites your change or conflicts, and the usual resolution is that somebody stops regenerating. Before adopting any tool, find out specifically how it handles handwritten additions, because every SDK eventually needs some.
Why do generated clients break when the API did not?
Because the SDK surface is derived from names in the document, not from the HTTP contract. Rename an operation identifier or restructure a schema and the HTTP behaviour is identical while every method and type name changes. This is why spec linting that pins operation identifiers matters so much: in an SDK-publishing organisation, those identifiers are public API even though they never appear on the wire.
Related reading
- Best OpenAPI and Swagger tooling — the spec hygiene that determines generated output quality in every tool here.
- Best API documentation tools — where generated snippets should come from so docs and SDKs cannot disagree.
- Best tools for API versioning and deprecation — the API-side commitment that keeps pinned SDK versions working.
- Best API testing tools — verifying that the client and the service still agree after a regeneration.
- Best CI/CD platforms — where the regeneration and multi-registry publish pipeline actually runs.
- Best API management platforms — how spec, SDK, docs and gateway decisions constrain one another.