Buyer’s Guide

Best Webhook Infrastructure Tools

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • webhooks
  • api
  • events
  • infrastructure

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

The webhook outage that teaches you the most is never a total outage. It is this one: a single customer deploys a change that makes their receiving endpoint hang for thirty seconds before timing out. Your delivery workers pick up their events, block, time out, schedule a retry, and pick up the next one. Within twenty minutes every worker in the pool is waiting on that one customer, and every other customer’s webhooks are hours behind. Nothing is down. Your dashboards are green. Support is not.

That is the defining property of webhook delivery and the reason it deserves dedicated infrastructure: you are making outbound HTTP requests to URLs that other people control, at a volume you do not control, with availability you cannot influence, and one bad consumer must not be able to degrade delivery for everyone else.

Most teams arrive here the same way. Webhooks started as a table and a cron job, which was completely correct at the time. Then a customer asked why they missed an event. Then someone asked for a way to replay. Then security asked why your service makes HTTP requests to arbitrary user-supplied URLs. Then a customer’s endpoint went down for a day and you discovered you had no circuit breaker. Each of those is a week, and none of them is the product you are trying to build.

This guide covers the four mechanisms that determine whether your webhook system is trustworthy, gives the honest version of the build-versus-buy argument including what “roll your own” actually costs over eighteen months, and then compares the platforms.

Key takeaways

  • At-least-once delivery is the only realistic guarantee, which means duplicates are certain. Your event IDs and your consumers’ idempotency handling are part of the contract, not an implementation detail.
  • Signatures must be computed over the raw request body. The most common integration bug in this category is a consumer framework parsing and re-serializing JSON before verification, which changes the bytes and fails every signature.
  • Ordering across a stream of webhooks cannot be guaranteed alongside retries and parallel delivery. Design events to carry a version so a consumer can discard stale ones, rather than promising order you cannot deliver.
  • Per-endpoint isolation is the feature that separates webhook infrastructure from a job queue. Without per-destination concurrency limits and circuit breaking, one slow consumer stalls every consumer.

The four mechanisms that make webhook delivery trustworthy

Retries, and what at-least-once actually commits you to

A webhook delivery attempt has three outcomes: a success status, an explicit failure status, or no answer at all. The third is the interesting one, because a request that times out may have been fully processed by the consumer before the connection dropped. You cannot distinguish that from a request that never arrived. Therefore you must retry, therefore duplicates are guaranteed, therefore at-least-once is the only honest guarantee to publish.

The retry policy itself is not complicated and is worth being explicit about:

  • Exponential backoff with jitter. Doubling intervals stop you hammering a struggling endpoint. Jitter stops every event queued during an outage from retrying simultaneously the moment the endpoint recovers, which is how you turn someone’s brief outage into a second outage caused by you.
  • A bounded attempt count with a dead letter destination. Retrying forever is a slow leak that eventually becomes a queue you cannot drain. Failed events should land somewhere durable, visible and replayable.
  • Retry only what is retryable. A 5xx or a timeout deserves a retry. A 400 means the consumer will reject it identically every time, and retrying is pure waste. A 410 should usually disable the endpoint. Treating all non-2xx responses the same is the laziest possible policy and it costs real capacity.
  • Replay as a first-class operation. When a consumer fixes their bug, the question is always “can you resend everything from Tuesday”. If the answer requires an engineer and a script, you will be asked every week.

Needs first-hand data: Pull a month of your own delivery attempts and classify every non-2xx response by status code, then report what share of total attempt volume was spent retrying responses that could never succeed, such as 400s and 422s. That share is wasted delivery capacity, the number is already sitting in your attempt log, and it is the cheapest argument you will ever make for a smarter retry policy.

What makes duplicates survivable is a stable event ID, sent in a header and in the body, that stays identical across every retry of the same event. Consumers key their idempotency on it. If your retries generate a fresh ID each time, you have made deduplication impossible and every consumer will process some events twice, and you will hear about it as a billing bug in their system rather than a delivery bug in yours.

Signature verification, and the bug everyone hits

Webhook authentication is the reverse of normal API authentication: the consumer needs to verify that a request genuinely came from you, without you having credentials on their system. The standard answer is an HMAC signature over the payload using a shared secret, sent in a header.

Four details decide whether it is actually secure:

Sign the raw body, verify the raw body. This is the failure that eats a day of every integration. A consumer’s web framework parses the JSON body into an object and hands the handler the parsed structure. The handler re-serializes it to compute the signature. Key order, whitespace and number formatting change, the bytes differ, and the signature never matches. The fix is always the same: capture the raw body before any parsing middleware touches it. Document this loudly, because your consumers will hit it.

Include a timestamp in the signed content. Signing the body alone means a captured request stays valid forever and can be replayed by anyone who observed it. Sign a timestamp together with the body, send the timestamp in the header, and have consumers reject anything outside a tolerance window. Tell them what window you recommend.

Support multiple active signing keys. Rotation with a single key requires simultaneous deploys on both sides, which means rotation never happens. Sign with the new key and the old one during an overlap window, send both signatures in the header, let consumers accept either, then retire the old key. A signing scheme without a rotation story is a scheme with permanent secrets.

Constant-time comparison. Comparing signature strings with a normal equality check leaks information through timing. Every language has a constant-time comparison function and your documentation should name it.

Mutual TLS is a stronger alternative some enterprise consumers will ask for, and it is a genuinely different operational commitment involving certificate distribution and rotation on their side. Offer it if your buyers are enterprises, do not make it the default, and keep HMAC as the path everyone else uses. The broader credential story is covered in the API authentication guide.

Ordering: stop promising it

Retries plus parallel delivery plus independent endpoints means events arrive out of order. That is not a defect you can configure away; it follows from the architecture. You have three options and only one of them is good.

Strict per-destination ordering. Single in-flight request per endpoint, next event blocked until the previous one is acknowledged. Ordering is guaranteed and throughput for that endpoint is capped at one over the round-trip time. Worse, a single slow event blocks the entire stream for that consumer, which is head-of-line blocking with a support ticket attached. Occasionally correct for genuinely sequential domains. Usually not worth it.

Sequence numbers with consumer-side reordering. You attach a monotonically increasing sequence per destination; consumers buffer and reorder. This pushes real complexity onto every consumer, including a decision about how long to wait for a missing sequence number before giving up. Most consumers will not implement it correctly, and you cannot tell which ones did.

Versioned state, which is the answer. Make the event carry the resource’s version or update timestamp along with the current state. A consumer receiving an update for a version older than what they already have discards it. Order stops mattering because the events are self-describing rather than incremental. This is the design that survives retries, duplicates, parallel delivery and consumer outages simultaneously.

There is a related decision worth making deliberately: whether the webhook carries the full state or only an identifier the consumer then fetches. Full state is fewer round trips and works when the consumer is down and replaying later, but it means whatever is in the payload has left your system, which is a data classification question. Identifier-only keeps payloads small and non-sensitive and guarantees the consumer sees current state, at the cost of an API call per event and a thundering herd against your own API whenever you fan out widely.

Per-endpoint failure isolation

This is the mechanism that distinguishes purpose-built webhook infrastructure from a general job queue, and it is the one that fails in the story at the top of this article.

A shared worker pool draining a single queue provides no isolation whatsoever. The slowest consumer consumes the most worker time, because slowness is measured in occupied worker seconds and a timeout is the most expensive possible outcome. The fix is a bulkhead: work for each destination is bounded so it cannot consume the shared resource.

Concretely, you need per-destination concurrency limits so no single endpoint can hold more than a set number of workers, aggressive and separately configured connect and read timeouts because a consumer that accepts your connection and never responds is worse than one that refuses it, a circuit breaker that stops attempting delivery after consecutive failures and probes periodically instead, and automatic endpoint disablement with notification after a sustained failure window so a customer who deleted their receiver stops costing you capacity indefinitely.

Per-endpoint rate limiting matters too, in the other direction: some consumers explicitly ask you to send no faster than a stated rate, and honouring that is a feature. It is the same token-bucket machinery discussed in the rate limiting guide, applied outbound instead of inbound.

Needs first-hand data: Inject a single destination that accepts connections and never responds into your delivery fleet, at a realistic share of total event volume, then measure end-to-end delivery latency for all other destinations over the following hour. Report the p50 and p99 shift. That one experiment tells you whether you have bulkheads or only believe you do, and it is a twenty-minute test in staging.

Rolling your own: the honest version

Do not let anyone tell you this is hard to start. A durable events table, a worker that selects due rows, an HTTP call, an attempt counter and an exponential backoff column is a day of work and it is genuinely correct for a first version. If you have ten consumers and they are all technical partners you talk to directly, that may remain correct for years. Buying a platform for that is overhead.

The honest question is what gets added afterwards, because the list is long and each item arrives as an urgent request rather than a planned project:

  • SSRF protection. Your service makes HTTP requests to URLs that customers supply. That is server-side request forgery by design. Without validation, a customer can point a webhook at your cloud metadata endpoint, an internal service, or a private address range, and use your own infrastructure to reach it. You need URL validation at registration, DNS resolution checks at request time to catch a hostname that resolves to a private address, redirect handling that revalidates each hop, and ideally delivery from an isolated egress path. This is the item most self-built systems are missing, and it is the one that becomes a security finding.
  • Static egress IPs. Enterprise consumers will ask to allowlist your source addresses. That means a stable, documented set of egress addresses, which means a NAT gateway or proxy layer and a commitment not to change them.
  • A customer-facing delivery log. Every event, every attempt, the response status, the response body excerpt, the timing. Without it, every “we did not receive the webhook” ticket is an engineer reading production logs. With it, most of those tickets never open, because the customer looks and finds their own 500.
  • Self-serve replay. Same argument. A customer who can replay a time range themselves does not file a ticket.
  • Endpoint ownership verification. A challenge-response at registration so somebody cannot register a URL they do not control and receive another tenant’s events.
  • Signing key rotation with overlap, per the section above, plus a UI for the customer to see and roll their own secret.
  • Payload size limits and truncation policy, because someone will eventually attach a large object to an event and your delivery pipeline will discover it.
  • Fan-out. One internal event, many subscribed endpoints per tenant, each with its own filters, retry state and circuit breaker. This is where a simple table schema stops being simple.
  • Documentation your consumers can implement against, including raw-body verification examples in several languages, which is the same production problem as the rest of your API documentation.

Add it up and the pattern becomes clear: the delivery mechanics are the small part. The product surface around them, meaning the log, the replay, the portal, the secret management, is the large part, and it is customer-facing, which means it needs design and support as well as code.

So the decision rule is not “is this hard”. It is: do your customers need to see and control webhook delivery themselves? If yes, you are building a product and buying is usually cheaper. If webhooks are an internal integration mechanism between systems you own, the table and the worker are fine and always were.

Svix

Svix homepage

Svix is webhook sending as a service. You send events to their API and they own fan-out, retries with backoff, signature generation and verification libraries, per-endpoint rate limits, and an embeddable consumer portal where your customers manage their own endpoints, view delivery logs and replay failures. That portal is the part that changes the economics, because it is the piece that takes support load off your team and it is the piece nobody wants to build.

Pros

  • The embeddable customer-facing portal gives your consumers delivery logs, endpoint management and self-serve replay without you building any of that UI
  • Signature scheme and verification libraries are published for many languages, so your consumers get working verification code rather than a paragraph of prose to implement
  • Per-endpoint rate limits, circuit breaking and automatic disablement with notification are built in, which is exactly the isolation layer that self-built systems lack
  • An open-source core exists, so a self-hosted deployment is possible if sending customer payloads to a third party is unacceptable

Cons

  • Every event payload passes through a third party, which is a data-flow question your security review will raise and which is unavoidable on the hosted path
  • It is outbound-focused, so webhooks you receive from other vendors are a different problem needing a different tool
  • Delivery is an availability dependency outside your control, and an incident on their side is one you explain to your customers without being able to fix
  • Self-hosting the open-source core means operating the delivery fleet and the datastore yourself, which recovers the residency answer and gives back the operational burden

Best for: Product teams whose customers need self-serve endpoint management, delivery logs and replay, and who do not want to build that portal.

Pricing: Usage-based metering on events sent and endpoints, with tiered plans layering on retention, throughput and enterprise controls, plus an open-source build you can self-host at infrastructure cost.

Hookdeck

Hookdeck homepage

Hookdeck treats webhooks as a gateway problem in both directions. Inbound, it receives webhooks on your behalf, queues them durably, filters and transforms them, and delivers them to your services with retries, which decouples your uptime from your vendors’ delivery schedules. Outbound, it handles sending. The inbound half is the distinctive part, because receiving webhooks reliably is a problem most teams solve badly with a synchronous handler that does real work before returning 200.

Pros

  • Inbound ingestion with durable queuing means a deploy or an outage on your side does not lose events from vendors whose retry policies are stingy
  • Filtering and transformation at the gateway keeps vendor-specific payload shapes out of your application code
  • Local development forwarding lets you receive real webhooks against a laptop without exposing a tunnel, which removes a genuinely annoying part of the workflow
  • Delivery logs and replay cover both directions, so one tool answers both “did we receive it” and “did we send it”

Cons

  • Sitting in the path for inbound events makes it a hard availability dependency for integrations you do not otherwise control
  • Managed-first, so a hard data residency requirement narrows your options quickly
  • Breadth across inbound and outbound means neither side is quite as deep as a specialist, and the outbound consumer portal is thinner than a dedicated sending platform
  • Transformation logic living in the gateway is code outside your repository, which is convenient until you need to review, test or roll it back

Best for: Teams whose bigger problem is reliably consuming webhooks from many third-party vendors, with outbound sending as a secondary need.

Pricing: Usage-based metering on events processed with tiered plans adding retention, throughput and team features.

Convoy

Convoy homepage

Convoy is an open-source webhook gateway you run yourself, covering both sending and receiving, with fan-out, per-endpoint rate limiting, circuit breaking, retry configuration, delivery logs and a portal you can expose to your consumers. It is the natural choice when the customer-facing feature set is what you want but the payloads cannot leave your infrastructure.

Pros

  • Self-hosted by design, so customer event payloads never leave your network and the compliance conversation ends before it starts
  • Covers the feature set that takes self-built systems eighteen months to reach: fan-out, per-endpoint limits, circuit breaking, replay and a consumer portal
  • Handles both inbound ingestion and outbound delivery, so one deployment covers both directions
  • Open source with no per-event licence cost, which changes the economics for high-volume senders where usage-based pricing dominates

Cons

  • You are now operating a stateful delivery system with a queue and a datastore, including its scaling, upgrades and on-call, which is the cost the managed options exist to remove
  • A smaller community than the hosted incumbents means fewer people have already solved your specific problem
  • The consumer-facing portal is functional rather than polished, and it is the surface your customers see
  • Self-hosting reintroduces the egress problem: static source addresses, SSRF controls and network isolation are yours to configure correctly

Best for: Teams that need customer-facing webhook delivery with logs and replay but cannot send payloads through a third party.

Pricing: Open source with no licence cost when self-hosted, plus a managed cloud offering metered on event volume.

Inngest

Inngest is not primarily a webhook sender, and including it is deliberate, because a large share of webhook pain is actually on the receiving side. It is a durable execution platform: events trigger functions, functions are composed of steps whose results are checkpointed, and a failure resumes from the last completed step rather than the beginning. Applied to webhooks, it turns “receive event, do five things, one of which is flaky” from a reliability problem into a defined workflow.

Pros

  • Step-level durability means a webhook handler that fails partway resumes where it stopped rather than re-running side effects it already performed
  • Concurrency and throttling controls are per key, so you can bound work per tenant and get the isolation property that a shared worker pool lacks
  • Long-running and multi-step flows including waits and delays are expressible in ordinary function code, without a separate workflow definition language
  • Retries, backoff and failure handling are platform behaviour rather than something each handler implements again

Cons

  • It is a processing platform, not a webhook sending product: no consumer portal, no per-customer endpoint management, no signature scheme for outbound delivery
  • Adopting it reshapes how your background work is written, which is a much larger change than adding a delivery service
  • Function code executing under someone else’s orchestration is a real coupling, and the local development and testing story requires investment
  • Pricing that meters steps rather than events can surprise you, because a workflow with many small steps is not obviously more expensive until it is

Best for: Teams whose real problem is reliably processing received webhooks through multi-step work, rather than delivering webhooks to customers.

Pricing: Usage-based metering on executed steps and concurrency with tiered plans, plus enterprise agreements above that.

QStash

QStash is a message queue with an HTTP interface: you publish a message addressed to a URL, and it handles delivery with retries, scheduling, delays and a dead letter queue. It is a primitive rather than a platform, and its value is that it needs no persistent connection and no worker process, which makes it work in serverless environments where a long-lived consumer is not an option.

Pros

  • Delivery, retry and scheduling with no queue infrastructure and no worker processes, which fits serverless and edge runtimes where a traditional consumer cannot run
  • The API surface is small enough to adopt in an afternoon, which makes it a genuine option for the case where a self-built table and cron is the alternative
  • Delays, schedules and dead letter handling are included, covering the scheduling half of what most teams build alongside webhooks
  • Signed delivery is supported so receivers can verify messages came from the expected source

Cons

  • No customer-facing portal, no per-endpoint circuit breaking and no consumer-visible delivery log, so it does not address the support-load problem at all
  • A general-purpose queue rather than webhook infrastructure, which means fan-out, subscription management and per-tenant endpoint configuration are yours to build on top
  • Delivery through a hosted service you do not control, with the same availability dependency as any managed option and less category-specific tooling around it
  • Positioned inside a broader serverless data platform, so evaluating it tends to pull in a wider vendor relationship than the single component you wanted

Best for: Serverless applications that need reliable delivery with retries and scheduling as a primitive, without a customer-facing webhook product around it.

Pricing: Usage-based metering on messages published with tiered plans by throughput and retention.

AWS EventBridge

EventBridge is an event bus with rules that route events to targets, and its API Destinations feature can deliver to external HTTP endpoints with managed connection credentials, retry policy, rate limiting and a dead letter queue. For teams already deep in AWS, it means outbound delivery lands inside the account, the IAM model and the bill you already have.

Pros

  • No new vendor, no new data processing agreement and no new on-call surface, which is frequently the strongest available argument in a regulated environment
  • Archive and replay of events on the bus are built in, so recovering from a consumer outage is a supported operation rather than a script
  • API Destinations handle credential management and invocation rate limiting per destination, providing the throttling half of endpoint isolation
  • Integrates directly with the rest of your AWS event sources, so internal events and outbound webhooks share one routing layer

Cons

  • Nothing here is customer-facing: no portal, no consumer-visible delivery log, no self-serve replay, so the support-load problem stays entirely with your team
  • Per-tenant endpoint management, subscription state and signing key rotation are all application code you write and operate on top
  • The signature scheme your consumers verify against is yours to design and implement, since the service authenticates outbound rather than signing for the receiver
  • Quotas, rule limits and payload constraints shape the architecture in ways that are painful to discover after you have built on it

Best for: AWS-centric teams delivering events to a bounded set of partner endpoints where no customer-facing webhook management surface is required.

Pricing: Usage-based metering on events published and on API Destination invocations, billed within the existing AWS account with no separate subscription.

How to choose

One question does most of the work: who needs to see the delivery log?

If the answer is your customers, you are building a product surface and the shortlist is Svix, Convoy or Hookdeck. Choose among them on data residency: Svix hosted if payloads may leave, Convoy self-hosted if they may not, Hookdeck if the inbound direction is the bigger pain.

If the answer is only your own engineers, you do not need webhook infrastructure at all. You need a durable queue with good retry semantics, which EventBridge or QStash provide inside an architecture you already run, or which your existing job system already does.

If the actual pain is processing events you receive rather than sending them, that is a different purchase and Inngest is the shape that fits it.

Two things to do before you commit, whichever way you go. First, run the slow-consumer experiment described above against whatever you have today, because the answer determines whether this is urgent or merely tidy. Second, write down your delivery guarantee in the words you would publish to customers: at-least-once, this retry schedule, ordering not guaranteed, deduplicate on this event ID header. If that paragraph is uncomfortable to write, the system is not ready regardless of which tool you pick.

OptionDirectionCustomer-facing portalSelf-hostPicks itself when
SvixOutboundYes, embeddableOpen-source coreCustomers manage their own endpoints and replay
HookdeckInbound and outboundPartialNoConsuming vendor webhooks reliably is the bigger problem
ConvoyInbound and outboundYesYes, by designYou need the portal but payloads cannot leave your network
InngestProcessing received eventsNoNoThe pain is multi-step work triggered by events
QStashOutbound primitiveNoNoServerless delivery with retries and no worker process
AWS EventBridgeOutbound to known partnersNoNoAlready on AWS and no customer-facing surface is needed

Frequently asked questions

Can I guarantee webhooks arrive in order?

Only by serialising delivery per destination and accepting head-of-line blocking, which means one slow event stalls that consumer’s entire stream. The better design is to include a version or update timestamp with the resource state in every event so a consumer can discard anything older than what it has already applied. Then ordering stops mattering, which is a stronger property than ordering.

How should consumers verify a webhook signature?

Capture the raw request body before any JSON parsing middleware touches it, compute an HMAC over the concatenation of the timestamp and that raw body using the shared secret, compare against the header value with a constant-time comparison, and reject anything whose timestamp falls outside a tolerance window. The raw-body step is where nearly every failed integration goes wrong, so put it first in your documentation with code in several languages.

Do I need webhook infrastructure if I only have a few integrations?

Probably not. A durable events table, a worker with exponential backoff and jitter, and a bounded attempt count is a correct first system and takes about a day. The trigger to buy is not volume, it is when your customers start needing to see delivery status and replay events themselves, because that is a product rather than a queue.

What is the security risk in letting customers register their own webhook URLs?

Server-side request forgery. Your infrastructure makes HTTP requests to addresses a customer chose, so without validation they can point you at internal services, private address ranges or a cloud metadata endpoint. Validate the URL at registration, resolve the hostname at request time and reject private addresses, revalidate on every redirect hop, and deliver from an isolated egress path. This is the control most self-built systems are missing.

Should webhook payloads contain the full resource or just an ID?

Full state with a version number is more resilient, because a consumer can process a backlog after an outage without hammering your API and without race conditions against current state. Identifier-only keeps sensitive data inside your boundary and guarantees freshness, at the cost of a fetch per event and a load spike on your own API during fan-out. If your payloads carry regulated data, identifier-only is the safer default.