A webhook receiver that doesn't double-charge treats every delivery as possibly a repeat. It verifies the signature, records the sender's event ID under a unique constraint before doing anything else, puts the event on a durable queue and returns a 2xx quickly. A separate worker then processes the queue idempotently and sends anything it can't handle to a dead-letter queue for inspection. Retries are not the bug; Stripe, GitHub, Microsoft Graph and Azure Event Grid all redeliver by design, and the receiver has to make the second delivery harmless.
Who this is for and what you will have
This guide is for developers and architects who receive payment, source control, Microsoft 365 or Azure events over HTTP and have seen duplicate orders, double emails or lost events. The examples use Azure Functions, Azure Service Bus and PostgreSQL, but the pattern carries over to any stack. At the end you will have:
- A table of how the common senders retry, time out and order events.
- A receiver that verifies, deduplicates and enqueues in a few milliseconds.
- Service Bus configured with duplicate detection and a dead-letter queue.
- An idempotent worker and a runbook for poisoned messages.
How webhook senders actually retry
Each sender has its own rules. Design for the strictest one you integrate with.
| Sender | Response deadline | Retry behaviour | Ordering |
|---|---|---|---|
| Stripe | Return 2xx quickly, before complex logic | Up to three days with exponential backoff in live mode; three times over a few hours in a sandbox | Not guaranteed |
| GitHub | 2xx within 10 seconds | No automatic redelivery; you redeliver failed deliveries manually or by code | Not documented |
| Microsoft Graph change notifications | 2xx within 3 seconds (10 seconds on retries) | Retries with exponential backoff for up to 4 hours; slow endpoints are delayed and then dropped | Not documented |
| Azure Event Grid | Waits 30 seconds | Schedule from 10 seconds up to every 12 hours, default 30 attempts and 1,440-minute TTL | Not guaranteed |
Details worth knowing:
- Event Grid treats only 200, 201, 202, 203 and 204 as success. It doesn't retry 400, 413, 401 or 403 from a webhook, and dead-lettering is off by default, so those events are dropped unless you configure a dead-letter storage container. Even when your endpoint recovers, Microsoft notes duplicates might still be received.
- Microsoft Graph marks an endpoint "slow" when more than 10% of responses exceed 3 seconds in a 10-minute window, and "drop" when more than 15% exceed the 10-second retry timeout. Dropped notifications can't be recovered.
- GitHub keeps the same
X-GitHub-Deliveryvalue when you redeliver, which makes it a natural deduplication key. - Stripe warns that the
createdtimestamp has one-second resolution, so it can't be used to order events or detect duplicates. Track event IDs instead.
The receiver architecture
Sender (Stripe / GitHub / Graph / Event Grid)
| HTTPS POST
v
Azure Function (HTTP trigger) <- must answer within the sender's deadline
1. verify signature or clientState
2. INSERT event_id into webhook_inbox (unique) -- duplicate? return 200 now
3. send to Service Bus queue with MessageId = event_id
4. return 202
|
v
Service Bus queue (duplicate detection on, max delivery count 10)
| peek-lock
v
Azure Function (Service Bus trigger) <- idempotent business logic
| failure after max deliveries
v
$deadletterqueue -> inspect, fix, resubmitThe inbox table and the queue each catch a different class of duplicate. The table catches the sender retrying a delivery you already accepted. Service Bus duplicate detection catches your own receiver resending after a crash between the queue send and the response. The idempotent worker catches redelivery of a message whose lock expired mid-processing.
Prerequisites
- An Azure Functions app (Node.js v4 programming model in the examples).
- A Service Bus namespace on the Standard or Premium tier. The Basic tier doesn't support duplicate detection.
- A database with unique constraints. The examples use PostgreSQL; if you run your own cluster, the Patroni high-availability guide covers keeping it available.
- The signing secret or client state value for each sender.
Step 1: Verify the request before trusting anything in it
Without verification, anyone can post a fake "payment succeeded" event.
- Stripe signs each delivery in the
Stripe-Signatureheader with HMAC SHA-256 over the timestamp and the raw body. Its libraries reject timestamps more than 5 minutes old by default; Stripe warns against a tolerance of 0, which disables the check. The raw body must reach the verifier unmodified, so disable framework body parsing on that route. - GitHub sends
X-Hub-Signature-256, an HMAC hex digest prefixed withsha256=. Compare with a constant-time function such ascrypto.timingSafeEqual, never==. - Microsoft Graph includes the
clientStateyou set on the subscription in every notification. Reject any notification whose value doesn't match.
Step 2: Choose the idempotency key
Use the sender's identifier, not a hash you compute:
| Sender | Key |
|---|---|
| Stripe | Event id; for fulfilment also data.object.id plus type |
| GitHub | X-GitHub-Delivery header |
| Event Grid | Event id |
| Microsoft Graph | Notifications carry an id, but the documentation doesn't present it as a deduplication key; combine it with subscription ID, resource and change type, and make the worker idempotent |
Stripe explicitly notes that in some cases two separate Event objects are generated for the same change, which is why the object ID and type matter for actions such as shipping an order.
Step 3: Record the key atomically
A unique constraint does the deduplication; application-level "check then insert" code races under concurrent retries.
CREATE TABLE webhook_inbox (
source text NOT NULL,
event_id text NOT NULL,
received_at timestamptz NOT NULL DEFAULT now(),
status text NOT NULL DEFAULT 'queued',
PRIMARY KEY (source, event_id)
);
INSERT INTO webhook_inbox (source, event_id)
VALUES ($1, $2)
ON CONFLICT DO NOTHING
RETURNING event_id;PostgreSQL returns rows from RETURNING only for rows actually inserted, so an empty result means the event was already accepted and the receiver can return 200 immediately without enqueuing again.
Step 4: Create the queue with duplicate detection
Duplicate detection can only be turned on when the queue or topic is created; you can't enable or disable it later, although you can change the window. The window defaults to 10 minutes, with a minimum of 20 seconds and a maximum of 7 days. Larger windows cost throughput, because every new MessageId is compared with the stored history.
az servicebus queue create \
--resource-group rg-integration \
--namespace-name sb-contoso-events \
--name webhook-events \
--enable-duplicate-detection true \
--duplicate-detection-history-time-window PT1HThe equivalent in Azure PowerShell:
New-AzServiceBusQueue -ResourceGroup rg-integration `
-NamespaceName sb-contoso-events `
-QueueName webhook-events `
-RequiresDuplicateDetection $True `
-DuplicateDetectionHistoryTimeWindow PT1HWhen a message arrives with a MessageId already logged in the window, Service Bus reports the send as successful and silently drops the new copy. If the queue is partitioned, uniqueness is MessageId plus PartitionKey, and Microsoft advises against combining deduplication and batching with partitioning.
Grant the receiver's managed identity the Azure Service Bus Data Sender role rather than using the namespace connection string. The worker needs at least Azure Service Bus Data Receiver; Microsoft's Functions documentation recommends Azure Service Bus Data Owner for identity-based trigger connections when you rely on accurate scaling, because without it the extension falls back to less accurate message estimation.
Step 5: The HTTP receiver
// src/functions/githubWebhook.js (Azure Functions, Node.js v4 model)
const { app } = require('@azure/functions');
const crypto = require('node:crypto');
const { ServiceBusClient } = require('@azure/service-bus');
const { DefaultAzureCredential } = require('@azure/identity');
const { recordInbox } = require('../inbox'); // runs the INSERT ... ON CONFLICT from step 3
const sb = new ServiceBusClient('sb-contoso-events.servicebus.windows.net', new DefaultAzureCredential());
const sender = sb.createSender('webhook-events');
function validSignature(rawBody, header, secret) {
if (!header || !header.startsWith('sha256=')) return false;
const expected = 'sha256=' + crypto.createHmac('sha256', secret).update(rawBody, 'utf8').digest('hex');
const a = Buffer.from(expected);
const b = Buffer.from(header);
return a.length === b.length && crypto.timingSafeEqual(a, b);
}
app.http('githubWebhook', {
methods: ['POST'],
authLevel: 'anonymous',
handler: async (request, context) => {
const rawBody = await request.text();
if (!validSignature(rawBody, request.headers.get('x-hub-signature-256'), process.env.GITHUB_WEBHOOK_SECRET)) {
return { status: 401 };
}
const deliveryId = request.headers.get('x-github-delivery');
const isNew = await recordInbox('github', deliveryId);
if (!isNew) {
context.log(`Duplicate delivery ${deliveryId} acknowledged`);
return { status: 200 };
}
await sender.sendMessages({
messageId: deliveryId,
contentType: 'application/json',
body: JSON.parse(rawBody),
applicationProperties: { source: 'github', event: request.headers.get('x-github-event') }
});
return { status: 202 };
}
});If the function crashes after inserting into the inbox but before the send completes, the sender's retry will see the row and skip the enqueue. To close that gap, have a small timer job re-send any inbox rows still marked queued after a few minutes; duplicate detection makes that re-send safe because the MessageId is the same.
Keep this function fast. An HTTP-triggered function that doesn't finish within 230 seconds gets an HTTP 502 from the Azure Load Balancer, but your real budget is far smaller: 3 seconds for Microsoft Graph and 10 seconds for GitHub.
Step 6: Process idempotently from the queue
The Functions runtime receives Service Bus messages in peek-lock mode. By default it completes the message when the function succeeds and abandons it when the function throws. Each abandon or lock expiry increments the delivery count; when it exceeds the queue's maximum delivery count (default 10), Service Bus moves the message to the dead-letter queue with reason MaxDeliveryCountExceeded.
const { app } = require('@azure/functions');
app.serviceBusQueue('processWebhook', {
connection: 'ServiceBusConnection',
queueName: 'webhook-events',
handler: async (message, context) => {
const id = context.triggerMetadata.messageId;
context.log(`Processing ${id}, delivery ${context.triggerMetadata.deliveryCount}`);
// Make the side effect idempotent: the payment or order call carries the same key every time
await fulfilOrder(message, { idempotencyKey: `webhook-${id}` });
await markInboxDone('github', id);
}
});Where the side effect is a call to another API, pass the event ID through as that API's idempotency key. Stripe's API, for example, accepts an Idempotency-Key header on POST requests: for API v1 it saves the first response for a key and returns it on retries, compares parameters and errors if they differ, allows keys up to 255 characters, and may prune keys once they are at least 24 hours old. Where the side effect is your own database, guard it with the same unique-constraint technique as the inbox, or with a state check such as "update the order to paid only where status is pending".
Ordering is not guaranteed by Stripe or Event Grid. Write handlers that tolerate an invoice.paid arriving before invoice.created, typically by fetching the current object from the sender's API instead of trusting the order of events. If you need strict ordering per entity, Service Bus sessions are the tool; the event-driven architecture comparison covers when a log-based broker fits better.
Step 7: Operate the dead-letter queue
The dead-letter queue is a subqueue that you address as webhook-events/$deadletterqueue. Microsoft is clear that there's no automatic cleanup: messages stay until you receive and complete them. Each dead-lettered message carries a DeadLetterReason and description.
| DeadLetterReason | Cause | Action |
|---|---|---|
MaxDeliveryCountExceeded | Handler kept failing, or locks expired before settlement | Fix the bug or dependency, then resubmit |
TTLExpiredException | Message expired before processing | Raise TTL or scale the worker |
Your own code, such as SchemaValidationFailed | Application dead-lettered a malformed payload | Correct the data or discard deliberately |
Dead-letter malformed payloads explicitly from code rather than throwing, so they don't burn ten delivery attempts. After fixing the cause, resubmit from Service Bus Explorer in the Azure portal, which lets you peek dead-lettered messages, edit them if needed and resend them individually or in batches. For Event Grid subscriptions, configure a dead-letter blob container, because otherwise undeliverable events are dropped.
Alert on dead-letter message count greater than zero. A dead-letter queue nobody watches is just a slower way to lose events.
Verify the design
- Send the same signed test payload twice. The first call returns 202, the second 200, and only one message reaches the queue.
- With Stripe, resend a real event using
stripe events resend <event_id> --webhook-endpoint=<endpoint_id>and confirm no second side effect. - Make the worker throw for one test message and confirm it lands in
$deadletterqueuewithMaxDeliveryCountExceededafter the configured number of attempts. - Kill the receiver between the inbox insert and the send, then let the timer job re-send. Confirm Service Bus drops nothing you need and processes the message once.
Troubleshooting
Stripe signature verification fails on every request. The framework parsed or re-serialised the JSON body. Stripe requires the raw body; any manipulation breaks verification.
Microsoft Graph notifications arrive late or stop. The endpoint is in the "slow" or "drop" state because responses exceed 3 seconds. Return 202 after persisting, or deliver Graph notifications to Event Hubs or Event Grid instead.
"Unauthorized access. 'Send' claim(s) are required to perform this operation." The identity sending to Service Bus lacks a data-plane role. Assign Azure Service Bus Data Sender on the namespace or queue.
Messages dead-letter with MaxDeliveryCountExceeded but logs show success. Settlement didn't reach the service, often because the lock was lost or the receiver closed first. Check lock duration against processing time.
Duplicates still reach the worker. The queue was created without duplicate detection, which can't be added later; create a new queue with it enabled. Also confirm messageId is set to the event ID, not a fresh GUID.
HTTP 502 from the function after a long wait. Work is running inside the HTTP trigger past 230 seconds. Move it to the queue worker.
Closing checklist
- Signature or
clientStateverified on the raw body before parsing. - Sender event ID stored under a unique constraint before any work.
- Service Bus queue on Standard or Premium, created with duplicate detection,
MessageIdset to the event ID. - Receiver returns 2xx inside the strictest sender deadline you support.
- Worker side effects keyed on the event ID; downstream APIs called with idempotency keys.
- Handlers tolerate out-of-order events.
- Dead-letter queue monitored, with a documented resubmit procedure.
- Event Grid subscriptions have a dead-letter container configured.
References
- Stripe: Receive Stripe events in your webhook endpoint
- Stripe: Idempotent requests
- GitHub: Best practices for using webhooks
- GitHub: Validating webhook deliveries
- GitHub: Handling failed webhook deliveries
- Microsoft Graph: Receive change notifications through webhooks
- Azure Event Grid delivery and retry
- Azure Service Bus duplicate message detection
- Enable duplicate message detection
- Service Bus dead-letter queues
- Azure Service Bus trigger for Azure Functions
- Azure Functions HTTP trigger
- Azure Functions Node.js developer guide
- Quickstart: Azure Service Bus queues with JavaScript
- ServiceBusMessage interface
- PostgreSQL INSERT