Authentication is a solved problem you can buy. Authorization is not, and the reason is that authorization is about your domain. Nobody else knows that a project belongs to a workspace, that a workspace has billing owners who are not administrators, that a document inherits permissions from its folder except when it has been explicitly shared, and that a support agent can read anything but only during an open ticket.
So every team writes it themselves, in if-statements, next to the handler. That works, until one day someone asks a question the code cannot answer: which users can see this document? Not “can this user see this document”, which is easy, but the inverse. And then a customer asks for custom roles. And then someone finds a place where the permission check was omitted entirely, in an endpoint written eighteen months ago by someone who has left.
The services in this post exist because that pattern is universal. They pull permission logic out of scattered conditionals into a place where it is defined once, testable, auditable and answerable in both directions. That is real value.
The cost is equally real and gets underweighted in every comparison: a permission check that used to be a comparison in local memory becomes a call to a service. It sits in the request path. It happens more than once per request. On a list endpoint it can happen once per row. Whatever that call costs, at the tail, is now part of what your users experience. That is the buying criterion, and most evaluations never measure it.
Key takeaways
- Model first, tool second. RBAC, ABAC and ReBAC answer different questions, and picking a tool before knowing which question you have is how teams end up modelling relationships as attributes.
- The decisive question is where the check executes: in your process, in a sidecar on loopback, or across the network. That choice determines your p99, not the vendor’s benchmark.
- List endpoints are where authorization services break. One check per row multiplies the tail latency of the authz call by the page size unless the product has a real filtering or lookup API.
- Google Zanzibar is a research paper, not a product you can buy. Several products implement its ideas, and they are not all the same.
RBAC, ABAC and ReBAC by worked example
Take one scenario and model it three ways. A document collaboration product. Documents live in folders, folders live in workspaces, users belong to workspaces, documents can be shared directly with individuals, and enterprise customers want their own role definitions.
RBAC: permissions attached to roles
Role-based access control assigns users to roles, and roles carry permissions. The check is a set membership test.
alice has role "editor" in workspace 42
role "editor" grants ["document.read", "document.write"]
can alice write document 7? -> is document 7 in workspace 42, and does alice have a role there granting document.write
This is the right model for most of what most applications need, and it is underrated because it is unfashionable. It is easy to explain to a customer, easy to display in an admin UI, and easy to store. If your permissions genuinely are “admins can do everything, members can do most things, viewers can read”, stop here. You do not need a service; you need a table and discipline.
RBAC breaks down in two specific places. The first is when roles need to be scoped to objects rather than to the whole tenant, so alice is an editor on this folder and a viewer on that one. Naive RBAC produces a role explosion, one role per object, and you now have a permissions table with a row per user per object which is exactly the thing you were trying to avoid. The second is when the decision depends on something other than identity, such as time of day, the resource’s state, or the request’s origin.
ABAC: permissions computed from attributes
Attribute-based access control evaluates a policy against attributes of the subject, the resource, the action and the environment.
permit if
subject.department == resource.department
and resource.classification <= subject.clearance
and action in ["read"]
and environment.time within subject.working_hours
and request.ip in corporate_ranges
This expresses things RBAC cannot. Contextual rules, data classification, regional restrictions, break-glass access with conditions. Policy is a function, and functions compose.
ABAC’s problem is the inverse query. Given the policy above, list every document alice can read. There is no index to consult. You either evaluate the policy against every document, which does not scale, or you write a second, hand-maintained query that approximates the policy and will eventually disagree with it. That divergence is a security bug with no test that catches it.
The other practical problem is attribute freshness. The policy needs the subject’s clearance and the resource’s classification at decision time. If those live in your database and the policy engine lives elsewhere, either you ship the attributes with every request or the engine fetches them, and one of those is a payload problem and the other is a latency problem.
ReBAC: permissions derived from a relationship graph
Relationship-based access control stores relationships between subjects and objects as tuples, and derives permissions by traversing them.
document:7#parent@folder:3
folder:3#parent@workspace:42
workspace:42#member@user:alice
document:7#viewer@user:bob
relation viewer on document = direct viewer or editor or (parent -> viewer on folder)
relation viewer on folder = direct viewer or (parent -> viewer on workspace)
relation viewer on workspace = member or admin
Now “can alice view document 7” is a graph traversal: document 7 has no direct grant for alice, so follow parent to folder 3, no grant, follow parent to workspace 42, alice is a member, permitted. And bob is permitted by a direct share without belonging to the workspace at all.
This is the model that fits collaboration products, hierarchical resources, and inherited permissions, which describes a large fraction of B2B SaaS. It also answers the inverse question well, because the graph is an index: “which documents can alice view” is a traversal from alice, and “who can view document 7” is a traversal from the document. That bidirectionality is the reason ReBAC systems exist.
The cost is that relationship data is a second source of truth. Every time your application creates a document, moves a folder, adds a member or revokes a share, it must write a tuple. If that write fails or is skipped in some code path, permissions silently diverge from reality, and the divergence is invisible until someone sees something they should not. Keeping tuples in sync with your domain data is the actual engineering work of adopting ReBAC, and it is far larger than writing the model.
How to pick between them
Most real systems need two of the three. The useful default:
- Use RBAC for coarse, tenant-wide capability: can this user administer billing, can they invite members.
- Use ReBAC for anything hierarchical or shareable: documents, folders, projects, nested organisations.
- Use ABAC for contextual constraints layered on top: time, location, resource state, data classification.
Most of the products below support more than one. What differs is which one is native and which one is emulated, and emulation shows up as awkward modelling and bad performance on exactly the queries you care about.
Zanzibar is a paper, not a product
You will see Zanzibar cited constantly in this category, so it is worth being precise. Zanzibar is a Google research paper describing the globally distributed authorization system behind their consumer products. It introduced the relationship-tuple model, the namespace configuration language, and a consistency mechanism using tokens, often called zookies, that let a caller demand a snapshot at least as fresh as a known write.
It is not software you can deploy. Google does not sell it. What exists are independent implementations of its ideas, with different storage backends, different consistency guarantees and different APIs. “Zanzibar-inspired” in a vendor’s marketing tells you the data model is relationship tuples. It tells you nothing about whether they implemented the consistency machinery, which is the hard and interesting part.
That machinery matters because of the problem the paper names directly: the new enemy problem. Remove someone from a document, then add sensitive content to it. If the permission check for that person reads a stale snapshot from before the revocation, they see the new content. The revocation happened first in real time and second in the replica the check consulted. Any authorization system with caching or replication has this problem, and the question to ask a vendor is not whether they are fast, it is what consistency options they offer and what the default is. A default of “eventually consistent, read from the nearest replica” is a correct choice for performance and a footgun if nobody told your engineers.
Where the check runs, and what it does to your p99
This is the buying criterion and it is almost never in the comparison tables.
An authorization check used to be a comparison against data already loaded in memory. When you adopt one of these services, it becomes one of three things.
In your process. An embedded library, or a policy engine compiled into or linked with your application, evaluating against data it already holds. No network hop. Latency is dominated by evaluation, not transport, and the tail is bounded by your own runtime. The tradeoff is data freshness: the process needs the policy and whatever data the policy consults, so something has to push updates to every instance, and every instance holds that state.
In a sidecar on loopback. A separate process on the same host or in the same pod, reached over localhost or a unix socket. You get process isolation and independent deployment without crossing the network. Tail latency is far better behaved than a remote call because there is no network variance, no cross-zone traffic and no shared connection pool contention. This is the shape most latency-sensitive deployments converge on.
Across the network. A central service, self-hosted or managed by the vendor. This is the model with the best consistency story, because there is one authoritative place, and the worst latency story, because every check crosses the network and possibly a region boundary. A managed authorization service in another cloud region sitting in your synchronous request path is a decision that deserves an explicit conversation rather than a default.
Then there is the part that turns a small cost into a large one: fan-out.
A typical request does not perform one authorization check. It checks the tenant, then the resource, then maybe a field-level rule. That is three sequential round trips before your handler runs. And a list endpoint returning fifty rows, each of which needs a visibility check, performs fifty. Sequentially, that is fifty round trips. In parallel, it is one round trip but fifty concurrent requests, and the endpoint’s latency is now the slowest of fifty samples from the authz service’s latency distribution. The tail of the slowest-of-N is much worse than the median of one, which is why an authorization service that looks fine in a single-check benchmark can devastate a list endpoint.
The fix is to not do it that way. Every serious product in this category has some form of bulk or filtering API: give me the subset of these IDs this user can see, or give me a filter expression I can push into my own database query. That capability is the single most important thing to evaluate, because it is the difference between authorization being a small constant cost and being proportional to your page size.
Needs first-hand data: Instrument your highest-traffic list endpoint and count how many authorization decisions one request actually requires. Then measure the endpoint’s p99 with the check stubbed out to return true, and again with the real service in the path, under production-like concurrency. The delta is what you are buying, and it is the only number in this evaluation that matters.
Needs first-hand data: Test the failure mode explicitly. Kill the authorization service or block it at the network level and see what your application does. Fail closed means an outage in the authz service is a full outage. Fail open means an authz outage is a security incident. Whichever you choose, verify it is what actually happens rather than what the configuration claims.
Decision logging is the other operational consideration. Authorization decisions are exactly the thing an auditor wants to see, and they are high volume: every check, with subject, resource, action, outcome and the rule that decided it. That is a genuine log volume problem and worth planning against your existing pipeline rather than discovering later, which is covered in log management tools.
OpenFGA

OpenFGA is an open-source relationship-based authorization engine derived from the Zanzibar model, developed under a neutral foundation and originating from Auth0’s fine-grained authorization work. You define a model in its own modelling language describing types, relations and how relations resolve through other relations, then write relationship tuples and ask check, expand and list queries. It is deployed as a service against a relational backing store, and it is the most direct open-source route to a real ReBAC implementation.
Pros
- A genuine ReBAC implementation with inheritance, indirect relations and both check and list-objects queries, not RBAC with relationship vocabulary
- Open source under neutral governance, which materially reduces the risk of adopting a critical-path dependency from a single vendor
- The modelling language is compact and reviewable, so permission semantics live in a file engineers can diff rather than in scattered conditionals
- List-objects and list-users queries answer the inverse question directly, which is the capability that saves you from hand-maintained approximation queries
Cons
- You run it, including the datastore, and it is now in the synchronous path of every authorized request
- Tuple synchronisation with your domain data is your problem entirely, and a missed write is a silent permissions bug with no natural test
- Consistency and caching behaviour need to be configured deliberately, and the convenient defaults are not the safe ones for revocation
- Purely relationship-oriented, so contextual and attribute-driven rules need to be expressed elsewhere or awkwardly encoded as relationships
Best for: Teams with hierarchical or shareable resources who need real relationship-based authorization, have platform capacity to run a stateful service, and want open governance rather than a single vendor.
Pricing: Open source with no licence cost. The cost is the infrastructure for the service and its datastore, plus the engineering time to keep tuples correct.
Oso

Oso started as an embeddable policy library with its own declarative language and has developed into a hosted authorization service alongside it. Its distinctive contribution is taking the list-filtering problem seriously: rather than asking the service about each row, it can produce a filter that is pushed down into your own database query, so authorization becomes part of the SQL rather than a fan-out of checks. For applications whose authorization pain is list endpoints rather than single checks, that is the right architectural answer.
Pros
- Data filtering pushes authorization into your own database query, which eliminates the per-row fan-out that ruins list endpoint latency
- The policy language expresses roles, relationships and attribute conditions together, so you are not forced to pick one paradigm and emulate the others
- Available as an embedded library and as a service, so the deployment shape can match your latency requirements rather than the vendor’s
- Policy as a reviewable file rather than console configuration, which keeps permission changes in the same workflow as code changes
Cons
- A domain-specific policy language is real learning cost, and the person who wrote the policy becomes a single point of knowledge unless you invest in that
- Data filtering needs an adapter for your data access layer, so ORM and database coverage is a hard constraint to check before committing
- The relationship model is less specialised than a dedicated Zanzibar-style engine for deeply nested hierarchies at large scale
- Commercial and open components have shifted over the product’s life, so pin down exactly what you are adopting and under what terms
Best for: Teams whose authorization problem is dominated by list endpoints and filtered queries, where pushing the decision into the database beats calling a service per row.
Pricing: A library with no licence cost alongside a commercial hosted service metered on usage, so the meter depends on which deployment shape you adopt.
Permit

Permit is a managed authorization platform built on open-source policy engines underneath, with a deliberate split: policy decisions are evaluated by a sidecar running next to your application, while the control plane, the policy authoring UI and the audit trail are hosted. That architecture is the interesting part, because it gives you local decision latency with a managed authoring and distribution experience. It also ships an end-user permissions UI, which addresses the requirement that appears when enterprise customers want to manage their own roles.
Pros
- Sidecar evaluation keeps the decision on loopback rather than crossing the network, which is the right answer for latency-sensitive request paths
- Supports RBAC, ABAC and ReBAC in one product, so a model that starts coarse and grows relationships does not require replatforming
- Policy authoring UI and embeddable end-user permission management cover the “our customers want custom roles” requirement without you building an admin surface
- Built on established open-source policy engines, so the evaluation semantics are not entirely proprietary
Cons
- The control plane is managed, so policy distribution depends on a vendor even though decisions are local, and you need to understand what happens when that link is down
- Running a sidecar per service instance is an operational pattern your deployment tooling must support, and it is real overhead in dense environments
- The abstraction over multiple underlying engines means debugging a surprising decision can require understanding the layer beneath
- Younger vendor for a critical-path dependency, which matters in enterprise procurement
Best for: Teams that want managed policy authoring and customer-facing role management without putting a network hop in every permission check.
Pricing: Metered on monthly active users and tiered by feature, with the sidecar deployment on your own infrastructure, so you pay for the control plane rather than per decision.
Cerbos

Cerbos is a stateless policy decision point you deploy yourself, as a sidecar or a service, that evaluates YAML-defined policies against a principal and a resource passed in with the request. The stateless design is the whole architecture: Cerbos holds no data about your users or resources, so there is no tuple store to synchronise and no second source of truth. You bring the attributes, it returns a decision. That makes it dramatically simpler to operate than a relationship-store system, and it draws a clear boundary around what it can answer.
Pros
- Stateless means no data synchronisation problem at all, which removes the single largest source of correctness bugs in ReBAC deployments
- Policies are YAML in version control, testable with the tooling it ships, so permission changes go through code review like anything else
- Deployable as a sidecar so decisions happen on loopback, and horizontally scalable trivially because there is no shared state
- Open source with an open core model, so an evaluation on the free build is broadly representative of the engineering reality
Cons
- Statelessness means you must pass all relevant attributes with every request, which pushes the data-fetching problem back into your application
- Answering the inverse question, meaning which resources a user can access, requires the query planning feature rather than being a natural property of a relationship store
- Deep hierarchical inheritance is expressible but is not the model’s centre of gravity the way it is in a Zanzibar-style engine
- Advanced management, distribution and audit capabilities sit in the commercial offering, so check the line before standardising on it
Best for: Teams that want policy decoupled from application code without adopting a second stateful system, and who can supply resource attributes at call time.
Pricing: Open source with no licence cost for the policy decision point, plus a commercial offering for managed policy distribution, audit and support.
Auth0 FGA

Auth0 FGA, also presented as Okta FGA, is the managed relationship-based authorization service from the same lineage that produced OpenFGA. The model language and query surface are familiar if you have looked at OpenFGA, with the operational burden removed and enterprise features layered on. For teams already inside the Auth0 or Okta relationship, it is the path of least procurement resistance, and it removes the stateful service you would otherwise run.
Pros
- Managed operation of a stateful authorization store, which is a genuinely unpleasant thing to run well and is the main reason teams hesitate on OpenFGA
- Shares its modelling approach with an open-source implementation, which gives you a more credible exit than a wholly proprietary engine
- Fits into an existing Auth0 or Okta commercial relationship, so procurement and security review are shorter
- Relationship model handles inherited and shared permissions natively, including the inverse queries that matter for list endpoints
Cons
- A managed service means every permission check is a network call to a vendor, which is the worst position for a check sitting in your synchronous request path
- Tuple synchronisation with your domain data is still entirely your responsibility, and being managed does not help with the hardest part
- Ties your authorization layer to the same vendor as your authentication, which concentrates risk and weakens your position at renewal
- Consistency and caching defaults need explicit attention, because a stale read is how revocations fail to take effect
Best for: Teams already on Auth0 or Okta who need relationship-based authorization and would rather not operate a stateful authorization store themselves.
Pricing: Metered on authorization checks and stored relationship tuples as part of the vendor’s platform pricing, so the meter grows with traffic rather than with user count.
Keycloak Authorization Services

Keycloak includes an authorization services layer built on the UMA specification, with resources, scopes, policies and permissions configurable in the same server that handles authentication. Policies can be role-based, group-based, time-based, or JavaScript you supply, and applications ask for decisions through a token exchange or an adapter. The appeal is obvious: if you already run Keycloak, you can define authorization in the same place as authentication without adopting another system.
Pros
- No additional service to deploy, operate or secure if Keycloak is already running, which is a significant practical advantage
- Policies and permissions sit alongside realms, roles and clients, so the whole identity picture is in one administrative surface
- Supports multiple policy types including attribute and time conditions, which covers contextual rules many RBAC systems cannot express
- No per-check or per-user cost, so authorization volume is not a budget conversation
Cons
- Decisions go through the Keycloak server, which makes your identity provider a synchronous dependency of every authorized request and a capacity planning problem you may not have signed up for
- Not a relationship-based model, so deeply nested inheritance and per-object sharing are awkward to express and worse to query
- Answering which resources a user can access is limited compared with purpose-built engines
- Configuration lives in the admin console and its export format rather than being naturally policy-as-code, which makes review and promotion between environments clunkier
Best for: Teams already running Keycloak whose authorization needs are role and attribute shaped rather than relationship shaped, and who would rather not add a second system.
Pricing: No licence cost, included in Keycloak. The cost is the additional load and capacity you are placing on the identity server.
How to choose
Model first. Write down five real permission questions from your product, including at least one inverse question of the form “list everything X can see”. If you cannot express them in your candidate’s model on paper, no benchmark matters.
| Your situation | The model you need | Where to start |
|---|---|---|
| Coarse tenant-wide roles, nothing nested | RBAC, probably no service at all | Your own database, or Keycloak Authorization Services if it is already running |
| Hierarchical resources, sharing, inheritance | ReBAC | OpenFGA self-hosted, Auth0 FGA managed |
| List endpoints dominate and per-row checks are the pain | Filter pushed into your query | Oso |
| Contextual rules, no second data store wanted | Stateless policy evaluation | Cerbos |
| Customers need to define their own roles | Managed authoring plus end-user UI | Permit |
| Latency budget is tight | Anything that evaluates locally or on loopback | Cerbos, Permit, Oso embedded |
Then run this sequence, in this order, because doing it backwards is how teams commit to a model that cannot answer their questions.
- Write the five permission questions. Include the inverse one. Express each in each candidate’s model.
- Count the authorization decisions in one request to your busiest list endpoint. That number is your fan-out.
- Decide where the decision executes. If the answer is a network call to another region, justify it explicitly or reject it.
- Work out who writes the relationship tuples or supplies the attributes, and where that code lives. This is the largest piece of work and it is always underestimated.
- Test revocation under realistic caching. Revoke access, immediately check, and confirm the answer. Then do it against a replica.
- Test the failure mode by breaking the dependency deliberately, and confirm the behaviour is the one you chose.
The relationship between roles in your product and groups coming from a customer’s directory is a separate problem worth reading about in SCIM provisioning tools, since group membership from an enterprise identity provider is where most role assignment originates in B2B products.
Needs first-hand data: Take your existing scattered permission checks and count them. Grep for the authorization helper across your codebase and count the call sites, then sample twenty and check whether any two implement the same rule differently. That inconsistency count is the honest argument for centralising, and it is far more persuasive internally than a vendor deck.
Frequently asked questions
Do I need an authorization service at all?
Probably not yet. If your model is genuinely roles per tenant, a table and a middleware function is correct and will stay correct for a long time. The signals that you have outgrown it are: you cannot answer “who can see this” without a full scan, permission logic has diverged between two endpoints, or a customer has asked for roles you did not define. Adopt when one of those is true, not in anticipation.
Should I put permissions in the JWT?
For coarse, slow-changing capability, yes, and it is the cheapest possible check. For anything fine-grained or frequently revoked, no, because a token is a cached snapshot and revocation does not reach it until the token expires. The practical split is coarse role claims in the token for routing and UI, authoritative checks server-side for anything that matters. The revocation tradeoff is the same one covered in session management libraries.
What is the difference between OpenFGA and Auth0 FGA?
OpenFGA is the open-source engine you deploy and operate. Auth0 FGA is the managed service from the same lineage. The model and query concepts are shared; the difference is who runs the stateful store, what the consistency and scaling story is, and whether every check crosses the network to a vendor. Choose on operational capacity and latency budget, not on the model.
How do I keep relationship tuples in sync with my database?
Write them in the same transaction boundary as the domain change wherever possible, or drive them from a change stream off your database so there is one path rather than many. The failure to avoid is scattering tuple writes across every handler that mutates data, because the handler someone adds next quarter will not include one. Add a reconciliation job that compares tuples against domain data and alerts on drift, and treat drift as a security alert.
Will an authorization service slow down my application?
Yes, by an amount that depends entirely on where it runs and how many times per request you call it. Local evaluation adds very little. A loopback sidecar adds a small, predictable amount. A network call to a managed service adds a round trip per check, multiplied by your fan-out, with the tail behaviour of whatever network sits between. Measure your own fan-out before believing any vendor’s single-check number.
Can I use one of these for API scopes as well?
You can, but scopes and application permissions are different granularities and conflating them tends to produce a model that serves neither well. Scopes limit what a client application may attempt; authorization decides what this user may do to this object. Keep them separate, and see API authentication tools and patterns for the client side.
Related reading
- Best authentication providers for developers — the layer underneath this one, and the build versus buy decision behind it.
- Best auth providers for B2B SaaS and enterprise — multi-tenancy models and where roles come from in a B2B product.
- Best SCIM provisioning tools — how group membership from a customer directory becomes role assignment in your product.
- Best open source authentication solutions — if self-hosting the authorization store, the same operational questions apply to the identity side.
- Best session management libraries — why putting fine-grained permissions in a token makes revocation a problem.
- Best API authentication tools and patterns — scopes, client credentials and the difference between them and user permissions.