Jump to content

Idempotency Beats Every Other Distributed Systems Guarantee You Are Paying For

Software Architecture Elena Rodriguez

What Happens When an Invoice-Paid Event Arrives Twice?

If processing the same invoice-paid event twice cannot grant a second entitlement, ordinary retryable delivery is often enough. That is the test worth running before anyone buys a more elaborate delivery guarantee.

Picture a billing webhook arriving while a tenant’s invoice is being marked paid. The handler verifies the tenant, checks the invoice, and grants access. Then the connection times out before the webhook provider receives a response. The provider sends the event again. From its side, the first attempt might have failed. From the application’s side, access might already exist.

The practical design has three parts: a key describing the business operation, a database uniqueness rule that enforces it, and a transaction that commits the entitlement alongside the outcome returned on replay. The aim is to prevent a duplicate business effect. Delivery can still happen more than once, and an email service or other external system does not become transactional just because the local database is tidy.

Put the Duplicate Check Where the Entitlement Is Written

Follow the event all the way through. The webhook provider sends a payload; the handler acknowledges it; the handler writes an entitlement row. A delivery claim made by the provider or broker describes the message path. It says nothing conclusive about whether that database row committed.

A provider event ID helps recognize a replay of the same event. It cannot necessarily recognize another event that requests the same grant. A billing operator might trigger a fresh invoice-paid notification with a fresh delivery ID while trying to repair an earlier failure. If the handler only deduplicates delivery IDs, it can grant the same access again.

Put the final guard on the authoritative entitlement write. For this invoice flow, a unique constraint spanning tenant, invoice, and entitlement operation can express the rule: this invoice grants this kind of access to this tenant once. The exact columns depend on what constitutes one grant in the product. The database can arbitrate competing inserts even when two handlers race; an application-level read followed by an insert cannot provide that protection on its own.

That distinction matters in software architecture reviews. A broker setting may sound reassuring, but the entitlement table is where double-provisioning becomes real.

Choose a Key That Means ‘This Entitlement, Once’

Suppose tenant Alder pays invoice INV-A for team access. A useful operation key combines the tenant ID, invoice ID, and entitlement type. The webhook’s event ID remains useful diagnostic context, but it does not define whether team access has already been granted for that invoice.

Operation version belongs in the key only when the domain permits another distinct grant for the same tenant, invoice, and type. If a changed payload merely describes the same grant differently, adding a version creates an escape hatch for duplicates. Decide what a second legitimate operation would mean before adding another field to the key.

Keep Payload Disagreements Visible

Store a fingerprint of the request fields that determine the entitlement. On replay, compare it with the stored fingerprint before returning the completed result. An identical request can receive the prior outcome. A request with the same business key but a different access scope must fail visibly. Quietly returning the old result would conceal a billing or mapping error; applying the new payload would let the caller rewrite an operation that was supposed to happen once.

Keep two retention questions separate. A short-lived dedupe record can help answer familiar webhook replays. The entitlement’s unique business rule must endure for as long as the entitlement remains authoritative, including after any broker cache or replay window expires. Deleting an old delivery record should never reopen the door to the same grant.

Commit the Dedupe Result and Entitlement Together

Start a database transaction after validating the request. Check the invoice state under the concurrency rules appropriate to that invoice, attempt the uniquely keyed entitlement insert, and record the completed outcome that a replay should receive. Commit those changes together. If the insert loses to another transaction, read the winning operation’s stored payload fingerprint and result rather than generating a fresh grant.

For a PostgreSQL implementation, PostgreSQL INSERT documentation describes the conflict-handling mechanism available at the write. The important property is database arbitration of the unique key. A separate “does this exist?” query can inform a response, but it must never serve as the only duplicate guard.

Image showing entitlement flow

When concurrent deliveries arrive, one transaction establishes the row. The other must wait for that transaction’s outcome or handle the resulting uniqueness conflict. If the first transaction rolls back, the second needs a path to succeed. A committed, permanent “processing” claim without a completed entitlement breaks that path and can suppress a valid retry.

If granting access should trigger downstream work, write an outbox record in the same transaction. That commits the entitlement and the intent to publish together. The outbox worker can send the message later. It cannot turn the eventual email delivery or remote API response into part of the database commit.

Commit The Outcome: A replay needs the result of the authoritative write, not a guess based on whether an earlier handler started running.

Give Replays a Domain Result, Not Another Side Effect

The handler contract should describe outcomes the business recognizes. A new invoice grant returns newly applied. A matching replay returns already applied with the stored result. The same key with a conflicting entitlement payload returns conflicting payload. A transient database failure returns a retryable failure.

That contract matters most around timeouts. A webhook provider can give up waiting while a transaction is still committing. Its next delivery cannot infer failure from the missing response. The handler must look up the business operation and return the committed outcome if it exists. For a recognized replay, a successful response with the prior payload also tells the provider to stop trying to repair an operation that already finished.

The outbox worker needs its own retry contract. It can claim pending work, attempt delivery, and resume after a crash. But consider the awkward moment after an email service accepts a welcome message and before the worker records success. A restarted worker may send that message again. If the destination supports an idempotency key, pass one derived from the outbox operation. If it does not, account for possible duplicates when choosing what the worker sends and how it records attempts.

“Exactly once” on a slide tends to skip that last sentence.

Know Which Guarantees a Dedupe Table Cannot Buy

A local unique constraint protects a local business invariant. It does not establish agreement among independently replicated stores, commit atomically across separate databases, or reverse an external action that has already happened. Those are different problems, with different failure points.

Ordered state changes deserve particular scrutiny. Suppose one service sees “invoice paid” while another processes a later reversal, each with its own database. Preventing duplicate grants in the entitlement table does not, by itself, establish which state transition should win across those services. The design may need versioned state transitions, a single authority for ordering, or explicit cross-service coordination. The requirement determines the mechanism.

Likewise, an outbox cannot make an email provider accept an operation only once unless that provider offers a suitable deduplication mechanism. Remote acceptance remains outside the local transaction. Treat claims of end-to-end exactly-once processing skeptically until someone names the consistency boundary and walks through a crash at every handoff.

Compare the operating burden of stronger coordination against the actual invariant and its failure modes. The cheap-looking architecture that loses paid entitlements is expensive. So is a coordination layer deployed to solve a duplicate that one authoritative unique index already prevents.

Pay for the Guarantee the Business Actually Needs

Take the invoice-paid flow through a concrete sequence before selecting infrastructure:

  1. Name the duplicate effect: the same invoice grants team access twice.
  2. Locate the authoritative write: the entitlement row, rather than the webhook receipt.
  3. Define the business key and the request fields whose disagreement must raise a conflict.
  4. Make the entitlement, reusable outcome, and any outbox intent commit atomically.
  5. Test simultaneous deliveries and a timeout after commit, then check what each retry returns.

Those tests tend to settle an argument faster than an industry opinion about messaging semantics. They also expose the cases worth escalating: perhaps access must follow a strict sequence across independent services, or a downstream destination cannot tolerate a repeated call. Bring in stronger coordination for the invariant that remains exposed, with that failure written down plainly.

Which specific business invariant still fails after the operation is made idempotent at its authoritative write?

Never Miss an Update

Fresh insights every week.

No spam. Unsubscribe anytime.

Your Thoughts

Share your thoughts.

Join the Discussion

Customise cookies