Before you add a retry to any background job, give the job an idempotency key for safe retries: one stable ID per intended effect, stored under a unique constraint and checked before the work runs. Then cap the attempts, about 5 as my starting rule, and move what still fails to a failed-jobs table you can replay from.

What an idempotency key for safe retries is

An idempotency key for safe retries is one stable ID per intended effect, derived from business facts such as the account and the billing period, and stored under a unique constraint. The worker records it with the effect, so a second attempt carrying the same key finds the record and stops instead of repeating the work.

An operation is idempotent when running it twice leaves the same result as running it once. RFC 9110 gives the HTTP version of that idea: a method is idempotent “if the intended effect on the server of multiple identical requests with that method is the same as the effect for a single such request”. Setting an order’s status to paid already behaves that way. Sending an email, charging a card, granting credits or creating a record does not, and making those effects safe to repeat is one of the controls in hardening a SaaS application for resilience.

The key is what makes such an effect safe. The worker writes the key together with the effect, or before it with a status that says whether the effect finished; on a second attempt, writing the same key again fails against the constraint, and the worker treats a finished key as a done job and moves on. The key travels on the job row or the queue message. It is enforced in one of two places: a processed-keys table, or a unique column on the effect’s own table, such as an idempotency_key column on a credit_grants table.

The key comes from the facts that define one intended effect, never from the attempt. These are my rules for the jobs a growing app runs most:

JobKey derived fromExample keyThe mistake
Welcome emailUser id, template name, template versionwelcome:{user_id}:{template}:{version}Keying on the email address, which changes when the user edits it
Monthly credit grantAccount id, billing periodcredits:{account_id}:{period}Using the time the job ran instead of the period it grants
Invoice sync to a CRMInvoice id, target systeminvoice-sync:{invoice_id}:{crm}Using the job’s own row id, which a second enqueue changes
AI generation billed per callA request id minted once, when the user clickedgenerate:{request_id}Minting the request id per attempt, so each retry is a new billed call

Two mistakes defeat the whole design: a new random key minted inside the retry loop, and a timestamp in the key. Both make every attempt look like a new effect, so the constraint never fires.

Idempotency key API design for your own endpoints applies the same rule at the HTTP layer: the client sends an Idempotency-Key header, and the server stores the key with the result of the first request. The header comes from the expired Idempotency-Key header draft at the IETF, whose page lists the document type as “Expired Internet-Draft” with revision 07 as the latest. In my reading, that leaves no standard behavior, so each API that uses the header sets its own rules.

Stripe’s idempotent requests reference is the working example. Stripe saves “the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails”, says keys can be removed once they are “at least 24 hours old”, and its idempotency layer “errors” when incoming parameters differ from the original request. All POST requests accept keys; sending one on a GET or DELETE request “has no effect”, because those requests “are idempotent by definition”.

When a job calls a provider that takes a key, the provider’s rules shape the key. Stripe suggests “V4 UUIDs, or another random string with enough entropy” and says to avoid “sensitive data (for example, email addresses or personal identifiers)” as keys. So mint a random UUID once, when the job is created, store it on the job row, and send that same value on every attempt. The business-fact key stays in your own table.

What is a dead letter queue, and the smallest version that works

A dead letter queue (DLQ) is a separate queue, or in a small app a table, that holds the messages a system could not process within the allowed attempts, so they stop being retried and wait for a person to inspect, fix and replay them. The smallest version is a failed-jobs table with the full payload and the last error.

AWS states the purpose in its SQS guide: dead-letter queues let you “isolate unconsumed messages to determine why processing did not succeed”. The dead letter queue meaning comes from the post office, where dead letter mail is, in Wikipedia’s words, “mail that cannot be delivered to the addressee or returned to the sender”.

A dead letter queue does not retry anything: nothing in it runs again until a person decides. A log is a poor substitute, because a log line rarely carries the payload a replay needs. And a dead letter queue nobody reads loses work as surely as a delete does.

My reading on when to skip one: work that is worthless late, such as a typing indicator or a cache warm-up, can be dropped and counted instead. AWS adds a caution for ordered work: “Don’t use a dead-letter queue with a FIFO queue if you don’t want to break the exact order of messages or operations”.

For an app whose queue is a Postgres table or a hosted runner, the smallest version that works is one more table, failed_jobs. The columns are my rules:

ColumnWhat it holdsWhy replay needs it
job_nameWhich handler ranThe replay sends the record back to the same handler
idempotency_keyThe key the job carriedA replay with the same key cannot double an effect that half-finished
payloadThe full job input, stored wholeThe job reruns without anyone guessing what it was asked to do
last_error, last_stackThe final error message and stack traceThe person fixing it starts from the failure, not from a reproduction
attemptsHow many tries the job usedShows whether the cap or a crash ended it
first_failed_at, last_failed_atWhen it first and last failedTies the record to a deploy or a provider outage
status, changed_byFailed, replayed or discarded, and who changed itNobody replays the same record twice, and every discard has a name on it

Two more rules go with the table. Personal data in a stored payload follows the same retention as the table it came from, and a non-empty failed_jobs table raises an alert. What each job’s log event should carry, the job id, the attempt number and the key, is part of error logging best practices for a solo app.

What goes wrong without them

Four failures follow from a retry with no key, or a last attempt with nowhere to land. The third column is my reading of the mechanism:

FailureWhat the user or the books seeWhich half prevents it
Duplicate effect: a call times out after the other side did the work, and the retry runs it againTwo welcome emails, two credit grants, a double chargeThe idempotency key
Silent drop: retries run out and the job disappearsThe receipt never sent, the sync never done, and nothing says soThe dead-letter path
Poison job: one bad payload fails on every attemptA worker’s time spent on the same failure at every retryThe attempt cap
Unreplayable failure: the only record is a console line without the payloadA failure you know about and cannot rerunThe stored payload

The duplicate happens because a timeout tells the caller nothing about the far side. The Idempotency-Key draft describes the case: after a request times out, “The client is left uncertain about the status of the resource”.

Take an app that sends each order’s receipt from a background job, retries when the email call fails, and deletes a job that runs out of attempts. One call times out after the provider has accepted the message, so the retry sends the receipt again; another job, whose payload the provider keeps rejecting, uses every attempt and is deleted with its payload, so nobody knows that customer never got a receipt. Both halves of the problem sit in that one case: a retry can duplicate an action, and a discarded failure can leave work unfinished. A key would have stopped the second receipt, and a failure record would have surfaced the missing one.

Across the 21 third-party apps I audited in June and July 2026, an 11-app deep-audit set and a 10-app held-out set, the Reliability & Correctness pillar averages 31.4 out of 100, the lowest of the 12 pillar averages. Those 21 are a selected set of audited apps, not a random sample, so the figure is no rate for AI-built apps in general. In the same audits, 17 of the 21 had no error tracking or alerting: when a user hits an error, nothing records it. My reading: in an app like that, a job that runs out of retries leaves no trace anyone sees.

The payment versions of these failures are six ways a vibe-coded checkout leaks money and a Stripe webhook that does not update the database. Events a provider sends you are the webhook side of the same control, which is how to make a webhook handler idempotent.

How to do it on Postgres, SQS, Azure and hosted runners

Reliable job execution takes 5 steps on any stack: derive the idempotency key, cap the attempts, write a failure record with the full payload, alert on the first record, and keep one command that replays a stored job with its original key.

The steps run in that order on every stack below; only the names of the settings change. Which errors deserve a retry, and how long to wait between attempts, is the subject of what exponential backoff is. Moving slow work off the request in the first place is how to run long tasks in the background.

A Postgres table as the queue: the key, the unique constraint and the failed-jobs table

Put a unique index on the job’s idempotency key column. Then enqueue with an insert that does nothing on a conflict: in PostgreSQL’s words, ON CONFLICT DO NOTHING “simply avoids inserting a row as its alternative action”, and with RETURNING, “Only rows that were successfully inserted or updated will be returned”. An empty result tells the caller the job already exists.

A key that no constraint enforces is a hope. I audited an AI coding workspace that saved every chat message through a database function, one copy of which was an upsert keyed on project and sequence number, but the table had no unique constraint on those columns; depending on which schema copy was live, every save errored, or retries and concurrent saves could write duplicate messages. The full telling is under a save path that assumed a uniqueness constraint.

When the effect lives in the same database, such as a row in a credit ledger, write the effect and the key in one transaction, so both land or neither does. When the effect is an outside call, my rule is key first: write the key to a processed-keys row with a status of pending, make the call, then mark the row done. A retry that finds pending knows an earlier attempt may have reached the provider, and sends the provider’s own idempotency key where the provider takes one.

Workers claim rows with FOR UPDATE SKIP LOCKED. PostgreSQL’s SELECT reference says it “can be used to avoid lock contention with multiple consumers accessing a queue-like table”, and warns that it is “not suitable for general purpose work”. Each claim adds one to attempts. Below the cap, a failure puts the row back to queued with its error; at the cap, the row is copied to failed_jobs and marked failed. A row left running by a crashed worker needs a sweep that returns it to queued after a timeout, or it never runs again.

-- 1. One job per intended effect
create unique index jobs_idempotency_key on jobs (idempotency_key);

-- 2. Enqueue: an empty result means the job already exists
insert into jobs (idempotency_key, name, payload)
values ($1, $2, $3)
on conflict (idempotency_key) do nothing
returning id;

-- 3. Claim: each worker takes a row no other worker holds
update jobs set status = 'running', attempts = attempts + 1
where id = (select id from jobs where status = 'queued'
            order by id limit 1 for update skip locked)
returning id, name, payload, attempts;

-- 4. At the cap: copy the whole job for replay, then mark it
insert into failed_jobs (job_name, idempotency_key, payload, last_error, attempts)
select name, idempotency_key, payload, last_error, attempts from jobs where id = $1;
update jobs set status = 'failed' where id = $1;

On Supabase the same pattern runs on the project’s own Postgres. Supabase Queues is the managed form, in Supabase’s words “a Postgres-native durable Message Queue system with guaranteed delivery built on the pgmq database extension”. Two open-source Postgres job libraries already do this work. pg-boss lists “dead letter queues with redrive, automatic retries with exponential backoff” and relies on Postgres’s SKIP LOCKED. Graphile Worker lists “Customizable retry count (default: 25 attempts over ~3 days)” and “Task de-duplication via unique job_key”. Whichever one runs the queue, the handler still needs its key, because a queue that hands a message out again has no way to know the email already went out.

Amazon SQS: the redrive policy, maxReceiveCount and redrive back to the source

SQS standard queues “ensure at-least-once message delivery”, and AWS’s SQS guide adds that “more than one copy of a message might be delivered”. So the handler needs its key whatever else you configure.

The dead-letter queue is a second queue, in the same AWS account and Region, named in the source queue’s redrive policy. The setting that moves messages there is maxReceiveCount, “the number of times a consumer can receive a message from a source queue before it is moved to a dead-letter queue”, and Amazon SQS dead-letter queues warns that with a value as low as 1, one failure to receive a message would move it there.

If a consumer does not delete a message before the SQS visibility timeout expires, the message “becomes visible again in the queue and can be retrieved by another consumer”; the default timeout is 30 seconds. My rule: set it above the job’s own timeout, so a slow attempt is not handed to a second worker while the first is still running. The per-attempt limit itself comes down to how to set a timeout on fetch on each outside call.

AWS calls it “a best practice to always set the retention period of a dead-letter queue to be longer than the retention period of the original queue”. Once the fix is deployed, dead-letter queue redrive by default “moves messages from a dead-letter queue to a source queue”.

AWS documents a CloudWatch alarm for messages moved to a dead-letter queue. Build it on ApproximateNumberOfMessagesVisible, the metric AWS recommends “To monitor the state of a DLQ”, because messages “automatically moved to a DLQ due to processing failures are not captured” by NumberOfMessagesSent. An alarm on that second metric stays quiet while jobs fail.

Windows Azure queues: Queue Storage and its poison-message queue

Windows Azure queues are today’s Azure Queue Storage, which has no automatic dead-lettering: each message carries a dequeue count, and the app decides what happens at the cap. An Azure Functions queue trigger does it for you, moving a message to a queue named with a -poison suffix after 5 failed attempts by default.

Windows Azure is the platform’s former name: Microsoft announced on March 25, 2014 that Windows Azure “will be renamed to Microsoft Azure, beginning April 3, 2014”. The Windows Azure queue service is now Azure Queue Storage, which Microsoft’s Azure Queue Storage introduction describes as “a service for storing large numbers of messages”. A message “can be up to 64 KB in size”, and when no time-to-live is set, “the default time-to-live is seven days”.

Queue Storage leaves dead-lettering to the app. Microsoft’s comparison of Storage queues and Service Bus lists automatic dead lettering as “No” for Storage queues and describes the pattern instead: the application examines each message’s DequeueCount, and past a threshold it moves the message to a dead letter queue the application defines.

The Azure Functions queue trigger does this for you. It “retries the function up to five times for a given queue message, including the first try”, then adds the message to a queue named <originalqueuename>-poison. The host.json settings for Functions queues set that count with maxDequeueCount, default 5: “The number of times to try processing a message before moving it to the poison queue.” My rule: the poison queue needs the same three things as any dead-letter path, an alert, an owner and a replay, because nothing reads it on its own. Azure Service Bus is the product with automatic dead-lettering; the differences that matter for this choice are in the questions at the end.

Hosted job runners and serverless functions

Hosted runners and queue-over-HTTP services retry for you and list failed runs in a dashboard, and in my reading that list is their dead-letter path. Two examples, in their docs’ words. In Inngest’s idempotency guide, a unique event id “acts as an idempotency key over a 24 hour period”, and a function’s idempotency expression “prevents another execution with the same value for 24 hours”. QStash dead letter queues work this way: QStash “automatically retries messages that fail due to a temporary issue but eventually stops and moves the message to a dead letter queue to be handled manually”, a message can be republished from the console, and the queue keeps messages for “a retention period that depends on your plan”.

Three things stay the app’s job on any runner, in my reading: the key derived from business facts and passed as the runner’s event id or idempotency field, the attempt cap, and an alert when a run lands in the failed list. Inngest’s guide gives the reason outside calls still need their own key: “A request can succeed before Inngest records the step result. A retry can then call the API again.” A scheduled job that fires twice is the same problem with the period in the key, whether it runs as a cron job every few minutes or once a month.

The recovery procedure: who is told, and how a job is replayed

These six steps are my procedure, written so a second person can follow them without the author on the call:

  1. 01 The alert fires on the first failure record and goes to the channel that pages someone, not a shared inbox.
  2. 02 Give the failure a severity by what the job does: a failed receipt is not a failed credit grant.
  3. 03 Read the error and the stored payload before changing anything.
  4. 04 Fix the cause, or mark the record discarded with a reason and your name.
  5. 05 Replay with one command or admin action that re-enqueues the stored payload with the same idempotency key, so a replay of work that half-finished cannot double it.
  6. 06 Confirm the effect happened once, then close the record as replayed.

For step 2, a failed credit grant can come close to what sev1 means for the accounts it touches; a failed receipt rarely does. On a Postgres queue, step 5 has one catch in my reading: where the jobs table itself holds the unique key, a fresh insert with the same key is skipped by the insert-or-skip, so the replay resets the existing row’s status and attempt count instead. Queue depth and the dead-letter count are reported by the worker’s own check, not by the web route’s health check endpoint.

How to verify it

Reliable job execution is proven with 4 tests: the same job enqueued twice produces one effect, two workers racing produce one effect, a job forced to fail uses every attempt and leaves a full failure record with an alert, and a replay from that record succeeds once.

Run them in staging against the real queue or runner, and keep the evidence each one names:

  1. 01 Replay: enqueue the same job twice with the same key and confirm one effect. Evidence: the two enqueue calls with their shared key, one entry in the provider's log (one email sent), one job row and one ledger entry.
  2. 02 Race: start two workers and give both the same job at once, as two enqueues with one key or one message both can claim, then confirm one effect and one clean skip. Evidence: the effect row and the second attempt's recorded skip, such as the empty insert result at enqueue, the second worker's key-exists log line, or the runner's record of a dropped duplicate.
  3. 03 Exhaust: make the job fail on purpose, with a flag or a payload the handler rejects, let it use every attempt, and confirm a failure record with payload, error and attempt count, and that the alert arrived. Evidence: the record and the alert with its time.
  4. 04 Recover: remove the fault, replay from the record with its original key, and confirm the effect happened once and the record closed. Evidence: the replay output and the closed record.

Keep each piece of evidence with its date, and rerun all four after a change of queue, runner or key format.

In the Production Hardening Sprint, deliverable 6.7 is verified this way: “Replay jobs, exhaust retries, and verify one intended result plus a recoverable failure record.”

Where the sprint does this

Deliverable 6.7, reliable job execution, is to “Make jobs idempotent and provide a dead-letter or failed-job path with a recovery procedure”, because “Retries can duplicate actions, while discarded failures leave work unfinished.” Webhook handlers are a separate deliverable, 5.2: “Make every webhook handler safe to repeat, including concurrent delivery and its downstream side effects.” Both land in the production readiness report, deliverable 13.1, which is verified this way: “Account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items.” Your app’s current framework and hosting setup are our starting point, and we refactor or replace components where the production work requires it. Hosting, paid tools, and API usage remain in your accounts. Each deliverable and its verify line is listed in the published scope.

Common questions about idempotent jobs and dead-letter queues

Does Kafka have a dead-letter queue?

Yes, in two of its components, though the broker’s own configuration lists no dead-letter setting. Kafka Connect’s sink connectors can write failed records to a dead letter queue topic named in errors.deadletterqueue.topic.name, which is blank by default, so nothing is recorded there until you set it (Kafka Connect’s error reporting). Kafka Streams has its own setting, errors.dead.letter.queue.topic.name: when it is not null, the default exception handler sends a dead letter queue record to that topic if an error occurs, and it is null by default. A consumer written without either one handles failed records in its own code, for example with a failed-records table of its own.

How does Kafka achieve idempotency?

Kafka achieves idempotency on the producer side with enable.idempotence: when it is true, “the producer will ensure that exactly one copy of each message is written in the stream”, and it is on by default “if no conflicting configurations are set”. My reading: that covers producer retries writing to the log, not the effect a consumer causes when it reads the message, which still needs its own key.

What are the key differences between Azure queues and Service Bus queues?

The key differences are dead-lettering, ordering, the delivery guarantee and the maximum message size. Microsoft’s Storage queues and Service Bus comparison lists automatic dead lettering as No for Storage queues and Yes for Service Bus queues. Storage queues give no ordering guarantee, while Service Bus gives first-in-first-out by using message sessions. Storage queues deliver at least once; Service Bus delivers at least once in its default PeekLock mode or at most once in ReceiveAndDelete mode. The maximum message size is 64 KB for Storage queues (48 KB with Base64 encoding) and 256 KB, 1 MB or 100 MB for Service Bus, depending on the service tier.

Which HTTP method is idempotent but not safe?

PUT and DELETE. RFC 9110 lists “PUT, DELETE, and safe request methods” as idempotent, and its safe methods are GET, HEAD, OPTIONS and TRACE. POST is in neither list, which is why a POST that creates something needs an idempotency key.