The identity provider is the one service where an outage is not degraded functionality, it is a closed business. Nobody logs in. Nobody’s token refreshes. Your own admin console is behind it, and so is the runbook you keep in a tool that uses SSO. That is a different operational category from the rest of your infrastructure, and it deserves to be designed as one.
Most guides to self-hosting auth compare feature matrices. The features are rarely what hurts. What hurts is the five things nobody demos: rotating a signing key without logging every user out, a session store that quietly becomes the highest-QPS service you own, a schema migration on a user table that is now large and cannot be rebuilt from anywhere, a disaster recovery plan whose restore has never been tested, and the discovery during an incident that your recovery path depends on the thing that is down.
This is written in the order you will meet those problems, not in the order a vendor would present them. The products come after, judged on how they behave when you are running them rather than on what they claim in a comparison table.
The prerequisite question, which is worth answering honestly before any of this: do you have someone whose job includes this system? Not someone who deployed it once. Someone who will be paged, will know the version, and will have practised the upgrade. If the answer is no, the right decision is a managed provider, and the full provider landscape covers those. Self-hosting identity without an owner is how you end up running an unpatched IdP for two years.
Key takeaways
- Key rotation only works without logging people out if relying parties re-fetch JWKS and you publish the new key before you sign with it. Overlap windows are the entire mechanism.
- Your session or token validation path will be higher QPS than your application. Size it independently and know exactly what happens when it fails.
- Migrations on the user table are one-way. Password hashes, MFA enrolments and consent records cannot be regenerated from anywhere, so the backup and the rehearsal are the plan.
- Design the IdP outage path before you need it: break-glass local admin access, and an on-call runbook that does not require logging in through the thing that is broken.
Start with the blast radius
Write down everything that stops working when the IdP is unavailable. The list is always longer than the mental model, and it changes the design.
Login stops, obviously. Existing sessions may continue if your applications validate tokens locally against a cached JWKS, which is the single strongest argument for stateless validation. If every request calls an introspection endpoint, then the IdP being down means every request fails, and you have converted an authentication outage into a total outage.
Token refresh stops. This is the delayed detonation. Users with valid access tokens keep working until those tokens expire, then they all fail at once. A short access token lifetime, chosen for good security reasons, sets the fuse length. Fifteen minutes of IdP downtime with fifteen-minute tokens is a rolling outage across your entire user base.
Machine-to-machine auth stops. Service-to-service calls using client credentials cannot get new tokens. Internal systems that looked unrelated to login begin failing, usually in an order nobody predicted.
Your own tools stop. If your monitoring, your CI, your cloud console and your incident tooling are behind the same SSO, the outage takes your response capability with it. This is the one that turns a thirty-minute fix into a three-hour one.
Two mitigations are worth building before you need them. First, a break-glass path: a local administrator credential for critical tools, stored offline, with its use alarmed. Second, longer refresh token lifetimes than instinct suggests, so a brief IdP outage does not cascade. Both feel like security compromises in a design review and both are what makes the system survivable.
Needs first-hand data: In staging, stop the IdP and record the exact minute at which each dependent system starts failing. Then stop it in the same way with your production token lifetimes configured. The gap between those timelines is your real recovery budget, and it is almost always shorter than the SLA you wrote down.
Signing key rotation without invalidating live sessions
This is the operation that teams postpone until a compliance requirement forces it, and then perform badly under time pressure.
The mechanism is straightforward once stated. Your IdP signs tokens with a private key. Relying parties validate them using the matching public key, fetched from a JWKS endpoint and cached. Rotation breaks when a relying party holds a cached key set that does not contain the key a token was signed with, and its behaviour in that situation is the whole ballgame.
The correct sequence has four steps and an overlap window in the middle.
One: publish the new key before you use it. Add the new key pair to the JWKS document while continuing to sign with the old one. Every relying party now has a chance to pick it up on its normal refresh cycle.
Two: wait longer than the longest JWKS cache TTL in your estate. This is the step people skip. The longest TTL is not the one you configured, it is the one in whatever library some team used three years ago with defaults nobody checked.
Three: switch signing to the new key, keep the old key published. Now tokens carry the new key ID. Tokens issued before the switch are still validatable, because the old public key is still in the document.
Four: remove the old key only after the longest possible token lifetime has passed. Refresh tokens live much longer than access tokens, so this window is measured by your longest-lived token, not your shortest.
Three things that break this in practice. A relying party that fetches JWKS once at startup and never again will fail at step three and keep failing until it restarts, and you will not know which services those are until it happens. A library that refetches JWKS on every unknown key ID is well behaved but becomes a thundering herd against your IdP at the moment of rotation, so rate limiting on that endpoint needs care. And a configuration that pins a single key rather than using discovery cannot rotate at all without a coordinated deploy.
Sessions themselves are a separate question. If your IdP keeps server-side sessions, rotating signing keys does not touch them. If sessions are represented purely by signed tokens, key removal is session termination, which is precisely why step four matters.
The related decision, stateless tokens versus server-side sessions, has consequences that run through everything below and is covered in session management libraries.
The session store becomes your highest-QPS service
Almost every team undersizes this, because the mental model is “login happens once a day per user” and the reality is that session or token validation happens on a large share of all requests.
Count the traffic honestly. Every page load that checks a session. Every API call that validates or introspects a token. Every silent refresh from a single-page application. Every mobile client waking up. Then add the multiplier nobody includes: a microservice architecture where five services each independently validate the same token on one user action.
Three architectures, and their scaling behaviour is genuinely different.
Server-side sessions in the primary database. Simplest to reason about and easiest to revoke, since deleting the row ends the session immediately. It also puts your highest-QPS read directly on the database that holds your user records, so a session read storm and a login are contending for the same resource. Works fine at moderate scale, fails in a way that takes user writes down with it.
Server-side sessions in a distributed cache. Moves the read load off the primary database and scales well. The new question is what happens when the cache is lost: if a cache restart logs out every user simultaneously, you have built an availability dependency on a system usually treated as disposable. Persistence or replication on that cache is not optional once it holds sessions.
Stateless tokens validated locally. The highest-QPS operation becomes a signature verification inside each service, with no network call at all. This scales essentially without limit and is the reason the pattern dominates. The cost is revocation: a token is valid until it expires, so logout and forced session termination are approximate rather than immediate, unless you add a revocation list, which is a shared lookup and quietly reintroduces the state you were avoiding.
The hybrid most mature systems land on: short-lived access tokens validated locally, long-lived refresh tokens stored server-side and revocable, with refresh token rotation so that reuse of an old refresh token is detectable and treated as theft. That gives you cheap validation with real revocation at the refresh boundary, and the worst-case exposure after a revocation is one access token lifetime.
Needs first-hand data: Instrument one week of production and record the ratio of token validations to logins. Then measure the p99 of a single validation under your session architecture, and again with the session store degraded to one node. Those two numbers size the store and tell you whether a cache failure is a slowdown or an outage.
Schema migrations on the user table are one-way
Ordinary application migrations are reversible in practice because you can rebuild the data. User table migrations are not, and it is worth being precise about why.
Password hashes cannot be regenerated. They are one-way by design. Lose them or corrupt them and every user resets their password, which is simultaneously a support event, a conversion event and a security event, because a mass password reset email is indistinguishable from a phishing campaign.
MFA enrolments cannot be regenerated. A TOTP secret lives in the user’s authenticator app and in your database. Lose your copy and every user re-enrols, through a recovery flow that is now handling your entire user base at once.
Consent and audit records cannot be regenerated. They are evidence, and their value is precisely that they were recorded at the time.
So the operating rules for anything touching this table are stricter than elsewhere.
Take the backup and verify it is restorable before the migration, not after. A backup you have not restored is a hypothesis.
Prefer expand-and-contract over in-place change. Add the new column, backfill it, write to both, switch reads, and only then drop the old one. Each step is individually reversible, which the single-step version is not.
Do not couple an IdP version upgrade with a schema change of your own. When something goes wrong you need to know which change caused it, and combining them removes that information exactly when you need it most.
Know how long the migration takes on production-sized data. A migration that takes eleven seconds on a development database can take hours on a large user table, and the difference is discovered during the maintenance window unless you rehearse against a restored production-sized copy.
Watch for locks on large tables. An operation that rewrites the table blocks writes, which means registration and password changes fail for the duration. Whether your database can do that operation online is the difference between a maintenance window and a maintenance weekend.
The specific trap in this category: the IdP’s own major version upgrade includes its own migrations, which you did not write and cannot easily inspect ahead of time. Run it against a restored copy of production first, every time, and time it.
Disaster recovery for the system that gates recovery
Standard DR planning assumes you can log in to execute the plan. For the IdP, that assumption is exactly what has failed.
Back up the database and the key material separately, and test both. Signing keys are often in a secret manager or a keystore rather than the database. A database restore without the matching keys produces a service that starts cleanly and rejects every existing token, which is a subtle and infuriating failure.
Keep the configuration in version control. Realms, clients, redirect URIs, identity provider connections, mappers, policies. Rebuilding those by hand from memory during an incident is how a recovery turns into a rewrite. Export and commit the configuration, and diff it periodically against what is actually running, because console changes drift.
Rehearse the restore, with a clock. Not “we have backups”. A timed exercise that produces a working login. The number that exercise gives you is your real RTO, and it is the number to put in front of anyone asking whether self-hosting was the right call.
Decide what happens to sessions after a restore. Restoring a database from six hours ago means sessions created since then are gone, and depending on your architecture the users holding those tokens either get logged out or, worse, hold tokens referencing users the restored database does not know about. Decide which of those you prefer and make it deliberate.
Have a break-glass administrator that does not depend on the IdP. A local account, credentials held offline, use alarmed and audited. Every organisation that has had a bad identity incident has this. Most that have not, do not.
Keep the runbook outside SSO. Print it, or put it somewhere that a person with no working login can reach. This sounds excessive until the first time it matters.
Needs first-hand data: Run a full restore drill from backup into an isolated environment: database, key material and configuration. Time it end to end, and confirm a real login works at the finish. Record what was missing on the first attempt, because something always is, and that list becomes the fix.
Keycloak

Keycloak is the most capable self-hosted IdP available and the most demanding to operate. It is a Java service, now a CNCF project, covering OIDC, OAuth 2.0, SAML and LDAP or Active Directory federation, with realms as the tenancy boundary and a clustering layer for distributed caching of sessions and configuration. Operationally the three things to internalise are JVM memory behaviour, the cache and clustering layer, and the fact that major upgrades have historically been significant events.
Pros
- Server-side sessions are first-class, so revocation and administrative session termination are immediate rather than approximate
- Key rotation is supported properly through the realm keys model, with old keys retained as passive for validation while the active key signs
- Configuration is exportable as JSON and therefore genuinely version-controllable, which makes disaster recovery a restore rather than a rebuild
- Every protocol you might need is present with no paid edition holding anything back, which removes a whole class of future surprise
Cons
- The distributed cache layer is the most common source of confusing production behaviour, and a partially split cluster produces inconsistent authentication results rather than a clean failure
- Memory sizing under load requires real attention, and the failure mode at scale is pauses and timeouts on the login path rather than an obvious crash
- Major version upgrades have changed the underlying platform and the admin API before, so upgrade planning is a project and rehearsal is mandatory
- Session-heavy deployments push serious load onto the database, so the session store scaling problem arrives earlier here than with token-only designs
Best for: Organisations with a platform team, a need for SAML and LDAP alongside OIDC, and a requirement that nothing be gated behind a licence.
Pricing: Free open source with no feature gating. The cost is infrastructure, a well-run database, and the permanent fraction of an engineer needed for upgrades, cache tuning and patch response.
Ory

Ory is a set of independent Go services rather than a single product: Kratos for identities and self-service flows, Hydra as an OAuth 2.0 and OIDC authorization server, Keto for permissions, Oathkeeper as an identity-aware proxy. Operationally this cuts both ways. Each service is small, stateless in the right places and easy to deploy, and you are running several of them with separate databases, separate migrations and separate upgrade cadences.
Pros
- Small stateless Go services scale horizontally without a clustering layer to reason about, which removes Keycloak’s most confusing failure mode
- You can adopt only what you need, so an OAuth authorization server does not oblige you to take on a user database and its migrations
- Database migrations are explicit CLI operations you run deliberately, which fits a controlled change process rather than surprising you at startup
- Hydra’s separation of the authorization server from the login and consent application means your UI can be restarted and redeployed without touching token issuance
Cons
- Multiple services means multiplied operational surface: several databases, several migration paths, several version compatibility matrices to keep aligned
- No admin console in the open source distribution, so routine administration is API calls and tooling you build
- The self-service flow model in Kratos is powerful and genuinely easy to implement incorrectly, and mistakes there are security-relevant rather than cosmetic
- Version compatibility between the components is a real constraint, and upgrading one without the others is how you find it
Best for: Teams with strong platform engineering who want composable, horizontally scalable identity services and intend to build their own interfaces and admin tooling.
Pricing: Open source under a permissive licence with no gated features in the self-hosted components, plus a separate managed cloud. Self-hosted cost is several services, their databases and the engineering time to build the surfaces they omit.
Authentik

Authentik is a Python-based IdP that has become the common answer for internal and infrastructure SSO. Architecturally it is a server plus worker processes plus outposts, which are separately deployed components that act as a reverse proxy or an LDAP endpoint in front of applications that cannot speak a modern protocol. The outposts are the distinguishing operational feature and the distinguishing operational responsibility.
Pros
- Proxy and LDAP outposts put SSO in front of legacy applications without modifying them, which no other project here does as cleanly
- Deployment is a normal container stack with a database and a cache, and it is markedly easier to stand up and reason about than a JVM cluster
- Flows and stages are configuration, so changing an authentication journey does not require a code deploy or a restart
- Configuration can be expressed as blueprints, which brings authentication configuration into the same review process as everything else
Cons
- Outposts are additional deployed components with their own health, versioning and network path, and they sit directly in the request path of the applications they front
- Open core, so confirm that the capability you are planning around is in the free build before you design against it
- The flow engine’s flexibility means a misordered stage can produce an authentication bypass that looks like a working configuration, and nothing warns you
- Smaller community than Keycloak, so unusual failures are more often yours to diagnose from source
Best for: Teams consolidating internal tool access behind SSO, especially where several of those tools have no OIDC support and need a proxy in front.
Pricing: Free open source core with a paid enterprise edition and support. Self-hosted cost is a modest container footprint, a database, a cache and the time to maintain flows and outposts.
Zitadel

Zitadel is a Go identity platform whose defining property is an event-sourced data model: state changes are appended as immutable events, and the queryable views are projections built from them. Operationally that is genuinely different from everything else here, and it needs to be understood before you scale it, not after. What it buys is an audit trail that is structurally complete rather than a feature someone remembered to enable.
Pros
- Complete immutable history of every identity change, which is the strongest audit and forensic position among self-hosted options
- Organisations and projects model multi-tenancy properly, including per-organisation identity providers and policies
- Same code runs self-hosted and in the vendor cloud, so moving in either direction later is a deployment change rather than a migration
- Passwordless and passkey flows are native, so the modern authentication path does not need bolting on
Cons
- The event store grows continuously and projection behaviour is unfamiliar, so capacity planning and performance debugging require learning a model most teams have not operated before
- Rebuilding projections after certain changes is a real operation with a real duration, and that duration on a large event stream is something to measure before you need it
- Smaller community than Keycloak, so operational answers come from documentation and the vendor rather than a decade of forum threads
- Mapping instances, organisations, projects and grants onto your tenancy design takes genuine thought, and correcting it later is expensive
Best for: Teams that need defensible audit evidence and real multi-tenancy, and are comfortable operating an event-sourced store.
Pricing: Free open source to self-host with a managed cloud metered on usage and commercial support available. Self-hosted cost is the service, its database and the learning curve on the storage model.
SuperTokens

SuperTokens is a lighter deployment than a full IdP: a core service that owns users and sessions, with SDKs in your frontend and backend that handle the rest. For self-hosting specifically that shape is attractive, because the operational surface is one stateless service plus your database, and the session design is the part of the product that has had the most attention.
Pros
- Small operational footprint: one service and a database you already run, with no clustering layer or cache coherence problem
- Session handling is the strongest part of the product, including refresh token rotation and detection of refresh token reuse as a theft signal
- Real server-side revocation, so logout and forced termination take effect immediately rather than at token expiry
- User data stays in your database, which resolves residency questions structurally rather than contractually
Cons
- The core service is on the login path, so a stateless service still needs redundancy and monitoring like anything critical
- Enterprise protocol support is thinner than the full IdPs, and SAML-shaped requirements will eventually push you elsewhere or into paid tiers
- Deep coupling to the SDKs means a future migration touches application code, not just configuration
- Less public operational experience at very high scale than the older projects here
Best for: Product teams self-hosting auth for one or a few applications who want correct session semantics without adopting a full identity platform.
Pricing: Open source core free to self-host with paid managed hosting and some advanced features on commercial tiers. Self-hosted cost is one small service plus existing database capacity.
FusionAuth

FusionAuth is a complete identity platform distributed as a single deployable, free to self-host, with one thing stated plainly: it is a commercial product with a free community edition rather than an OSI open source project. Operationally it is one of the easier full platforms to run, because it is a single application with a database rather than a cluster of components, and its migration tooling is unusually good.
Pros
- Single deployable plus a database is the smallest operational surface of any full-featured platform here
- Support for a wide range of legacy password hash formats, which lets you migrate an existing user population without forcing a reset
- Upgrade process is well documented and predictable, which for a system this central is worth more than most features
- Configuration and tenant setup are fully driven by an API, so environments can be rebuilt from code
Cons
- Not OSI open source, which blocks it outright under some procurement policies and constrains redistribution
- Meaningful capability including some enterprise protocol and threat detection features sits behind paid editions, so plan against the edition you will actually run
- Single vendor with no community fork, so the licence terms and the roadmap are not negotiable
- Configuration breadth means the admin console has a learning curve comparable to Keycloak’s
Best for: Teams moving off a home-grown auth system who need to import existing password hashes intact, and who can accept a source-available licence.
Pricing: Free community edition for self-hosting, with paid editions and hosted plans unlocking advanced features and support, metered on usage rather than a flat licence.
Logto

Logto is a newer self-hostable identity service built for developer experience: an OIDC provider, a configurable sign-in experience, a management console and a management API, with organisations for B2B tenancy. Operationally it sits at the light end, a service plus a database, and the honest assessment is that its maturity for large self-hosted deployments is still being established in public.
Pros
- Quick to a working, branded login without the deployment weight of a full platform
- Management API is well shaped, so creating tenants, applications and connections is automatable from the start
- Sign-in experience is configured rather than built, which removes the theme-development frustration of heavier platforms
- Modern SDK coverage means the application side of the integration is short
Cons
- Younger project, so long-running self-hosted deployments, upgrade history and large-scale behaviour have less public evidence behind them
- Open core, and the boundary sits close to features B2B products need, so confirm the free build covers your requirements
- Less extensible than Keycloak or Ory when custom logic has to run inside the authentication flow
- Operational documentation for scaled self-hosting is thinner than for the established projects
Best for: Small teams self-hosting identity for a product that fits the platform’s model, who value integration speed over extensibility.
Pricing: Open source core free to self-host, with a managed cloud and paid tiers gating some enterprise and organisation features.
How to choose
Pick by operational shape first, then by protocol requirements, then by everything else. That order is deliberate, because the shape is what you live with daily and the protocols are what block you occasionally.
If you need SAML and LDAP federation with nothing gated, it is Keycloak, and you should staff it accordingly. Nothing else in the free tier covers that ground as completely.
If you have strong platform engineering and want horizontal scaling without a cluster to reason about, Ory. Accept that you are building the admin tooling and the UI.
If the job is internal SSO including applications that cannot speak OIDC, Authentik, because the outposts solve that specific problem directly.
If audit evidence and multi-tenancy are the drivers, Zitadel, with time budgeted to learn the event-sourced model before you depend on it.
If it is one product and you mainly want sessions done correctly, SuperTokens is the smallest thing that is genuinely good at that.
If you are migrating existing password hashes, FusionAuth, provided a source-available licence passes your policy.
If you want a light, modern service and fit inside its model, Logto.
| Project | Deployment shape | Session model | Upgrade risk | Operational demand |
|---|---|---|---|---|
| Keycloak | JVM service plus clustered cache | Server-side, immediate revocation | High across majors | High |
| Ory | Several Go services, separate databases | Configurable per component | Medium, compatibility matrix | High |
| Authentik | Server, workers and outposts | Server-side | Medium | Medium |
| Zitadel | Go service on an event store | Server-side plus tokens | Medium | Medium, unfamiliar model |
| SuperTokens | One core service plus your database | Rotating refresh, real revocation | Low | Low |
| FusionAuth | Single deployable plus database | Server-side | Low, well documented | Low to medium |
| Logto | Single service plus database | Server-side plus tokens | Unproven at length | Low |
Whatever you pick, do four things in the first month: commit the configuration to version control, run a timed restore drill, rehearse a key rotation in staging, and write the break-glass procedure. Those four are what separate a self-hosted IdP from a self-hosted incident.
Frequently asked questions
How often should I rotate signing keys?
Frequently enough that the procedure is routine rather than exceptional, which for most teams means quarterly or twice a year. The specific interval matters far less than whether you have ever done it successfully. A team that rotates on a schedule discovers the badly behaved relying party during a planned window. A team that rotates for the first time during a suspected key compromise discovers it at the worst possible moment.
Should I use stateless JWTs or server-side sessions?
Both, in the arrangement most mature systems converge on: short-lived access tokens validated locally so validation costs nothing at scale, and long-lived refresh tokens held server-side so revocation is real. Refresh token rotation on top of that turns a stolen refresh token into a detectable event rather than a silent persistent compromise. Pure stateless gives you scale without revocation. Pure server-side gives you revocation with a scaling ceiling. The hybrid is the honest answer and the tradeoffs are laid out in session management libraries.
Can I run a self-hosted IdP in a single region?
Usually yes, and multi-region is a much bigger project than it first appears, because identity state is write-heavy and consistency-sensitive in ways that resist naive replication. A single region with a well-tested failover and a documented RTO is a defensible design for most organisations. What is not defensible is a single region with untested backups, because the failure mode there is a business-ending outage rather than an inconvenient one.
What breaks first when a self-hosted IdP is under-provisioned?
The login path, and it degrades rather than failing cleanly. Password hashing is deliberately CPU-expensive, so a login surge saturates CPU before anything else, and the visible symptom is slow logins and timeouts rather than errors. The second thing is database connections, since every login and every session read needs one. Alert on login latency percentiles and on connection pool saturation specifically, not on general host metrics, and keep those alerts in a channel that does not depend on the IdP. Capturing the authentication audit trail somewhere durable is a log management problem worth solving at the same time.
Is self-hosting cheaper than a managed identity provider?
At large user counts, clearly yes, because managed pricing is metered on users and self-hosting is metered on infrastructure. At small user counts, clearly no, because the engineer time exceeds any plausible bill. The losing case is the middle, where you pay a fraction of a salaried engineer to avoid a subscription that costs less. Work out your projected user count and your loaded engineer cost before assuming. The cost comparison across managed options is in best authentication providers.
Related reading
- Best authentication providers for developers — the managed side of this decision and the build-versus-buy arithmetic.
- Best open source authentication solutions — the same projects judged on licence boundaries and features rather than operations.
- Best session management libraries — token versus session design, revocation and refresh rotation in depth.
- Best CIAM platforms — what self-hosting has to add when the user population is consumers.
- Best enterprise SSO solutions — what your self-hosted IdP has to do once customers bring their own identity providers.
- Best SCIM provisioning tools — automated user lifecycle on top of the IdP you just stood up.