Idempotency Is a Client Contract, Not a Server Flag
Rowland Adimoha / September 26, 2026
33 min read
Rowland Adimoha / September 26, 2026
33 min read

Retrying a request that moves money is safe only when both ends of the connection agree on what counts as the same request. For an HTTP POST, that agreement is what idempotency means. The server cannot switch it on alone. It is a contract with two parties, and the client writes the first clause by choosing a key.
The distinction matters because a broken contract fails quietly. A duplicated charge throws no exception. Both requests return 201 Created, both ledger rows balance, and the error surfaces weeks later as a refund ticket or a reconciliation mismatch. Teams that treat idempotency as a server setting tend to discover the gap in finance's spreadsheet rather than in their own logs.
The explanation below builds the model in the order the problem forces on you. It starts with why retries cannot be avoided and who is able to know whether two requests share one intent. Then it turns to the server's side of the contract, which covers what the store holds, where the effect runs, how races resolve, which outcomes get replayed, how long records live, and how the server recovers from its own crashes. The last sections cover the gaps that still let duplicates through, how to test for them, and when a key is the wrong tool. Go and SQL snippets show the hinges. Animated scenes and a few diagrams carry the flows.
A client that sends a request can end up in one of three situations. The server processed the request and the response arrived. The server rejected or failed the request and that response arrived. Or nothing arrived before the client's deadline.
1. Worked, and you heard
2. Failed, and you heard
3. Nothing came back
Never ran, or ran and the reply was lost?
The third ending is the whole problem. A timeout does not tell the client which side of the network the failure happened on. The request may have died in a load balancer before any handler saw it. Or the handler may have written the ledger row and then lost the connection while writing the response. The client sees the same symptom in both cases, a socket that stopped producing bytes, and nothing on its side of the wire can tell the two apart.
This is the two generals problem in the form every payment engineer eventually meets. No finite exchange of messages over an unreliable link lets both parties be certain that the other knows the outcome. Adding an acknowledgment only moves the uncertainty onto the acknowledgment. Distributed systems do not solve this problem. They work around it by making the uncertain action safe to repeat.
After a timeout the client has two choices, and both are wrong without help. It can give up and tell the user the payment failed, which is false whenever the server did charge. Or it can retry, which is false whenever the first attempt succeeded, because the retry becomes a second payment. Real clients retry. Mobile SDKs retry on connection reset. Service meshes and load balancers retry on upstream failure. HTTP libraries retry when configured to, and some decide on their own, as the Go standard library does in a case covered below. Users retry with their thumbs.
The scene above plays one checkout on a loop. The first POST /v1/charges writes ch_8k1, the response is lost on the way back, and the user taps Pay again. The story ends with one charge instead of ch_8k2 only because the server could answer one question. Is this request the same as the first one? Everything else in this explanation is about who gets to answer that question and how the answer is kept.
The word arrives in design reviews carrying at least four meanings. They overlap enough to be confused and differ enough to cause bugs.
Mathematical. An operation is idempotent when applying it twice gives the same result as applying it once. Setting an order's status to paid is idempotent. Adding 4,200 to a balance is not.
HTTP method semantics. RFC 9110 calls a method idempotent when the intended effect on the server of several identical requests equals the effect of one. PUT, DELETE, and the safe methods GET, HEAD, OPTIONS, and TRACE qualify. POST and PATCH do not. The definition concerns effect, not response. A second DELETE /orders/42 may return 404 instead of 204 and still be idempotent, because the order is gone either way.
Deduplication. A server or consumer drops a request it recognises as already handled. This is a mechanism rather than a property, and it only works if each request carries an identity the receiver can recognise.
Exactly-once processing. An effect happens one time no matter how many times the message is delivered. This is an end-to-end outcome. A broker can provide it inside its own transactional boundary. It cannot provide it for an effect outside that boundary, such as an authorisation at a card network.
When someone asks to "make the charge endpoint idempotent," they want the fourth meaning and usually propose the second. POST /v1/charges is not idempotent by method, and no configuration changes that. What a team can build is deduplication keyed on an identity that represents the customer's intent. That produces exactly-once charging as a property of the application, which is the thing the business wanted all along. The rest of this piece is about that identity and the machinery that honours it.
The tempting shortcut is to let the server decide sameness from the request itself. Hash the body, remember the hash for a few minutes, and reject repeats. This fails in both directions.
Identical request bodies from a retry and from a genuine second order
Consider a café app. A customer buys a flat white for 350 pence, then turns to a friend and buys a second one forty seconds later. Both requests carry the same customer, card, amount, and currency. The bodies are byte for byte identical. A content hash treats the second order as a retry and drops it. The café loses a sale, and the customer wonders why the receipt shows one coffee.
Now reverse it. A genuine retry passes through middleware that adds a request_timestamp field, or through a newer SDK that serialises JSON keys in a different order. The bytes differ, the hash differs, and the server charges twice. Content deduplication blocks legitimate repeats and lets real retries through, which is the worst combination available.
The cause is that intent is not in the bytes. Two purchases can produce identical payloads, and one purchase can produce different payloads across attempts. Only the party that decided to make the purchase knows which is which.
The same limit applies to a framework flag. A route annotated idempotent: true has no more information than the content hash does. It sees two requests arrive. It cannot know whether a person meant one payment or two. At best the flag switches on deduplication keyed on something, and if that something is not an identity the client chose, the flag inherits every failure above.
A claim that an API is idempotent can be tested with four questions. By which key? Scoped to what? Stored where? Kept for how long? If the answers are missing, the claim means that retries are rare and the team hopes they stay that way.
Sameness lives with whoever made the decision, so the client has to name it. The Idempotency-Key request header is the conventional place. The IETF httpapi working group's draft for the header says a key must be unique and must not be reused with a different request payload, and recommends a UUID or similar random identifier. Stripe accepts keys up to 255 characters, suggests version 4 UUIDs or other high-entropy strings, and warns against putting personal data such as email addresses in them.
The client's rules are few and strict.
Mint the key at the moment of commitment. That moment is when the user presses Pay, or when a job decides to issue a refund. It is not when the HTTP library builds a request. A key minted per attempt makes every retry look new, and the server will correctly treat each one as a fresh payment.
Persist the key with the intent, not with the attempt. On mobile, store it alongside the checkout session in local storage, so an app killed mid-request resumes with the same key after relaunch. On the web, keep it in the checkout session on your own backend rather than in a JavaScript variable that dies with the tab.
Reuse it on every automatic retry of that intent. Backoff loops, connection-reset retries, and resumes after a restart all send the same value.
Mint a new key for a new intent. A different amount, a new cart, another card after a decline, or a user who deliberately starts over is a new intent.
Never change the body under an existing key. If the amount changed, the intent changed, and the key changes with it.
Callers other than people need keys too, and the right source differs by caller. A browser or app checkout mints the key at the Pay tap and stores it with the session. A queue consumer derives its key from a business identifier in the payload, such as order_7731:capture, because that identifier survives redelivery and republishing. A client that creates a resource under an identifier it chose, as in PUT /orders/7731, already holds a natural key and may not need the header at all. The test for any source is the same. Every retry of one intent must produce the same value, and a new intent must produce a new one.
One detail in the Go standard library shows why servers cannot treat the header as decoration. The net/http transport retries a request after certain network errors on a reused connection only when it considers the request idempotent. It treats GET, HEAD, OPTIONS, and TRACE that way. It also treats any request whose headers contain Idempotency-Key or X-Idempotency-Key as idempotent, whatever the method. Adding the header to a POST therefore changes client behaviour on its own, because the transport may now resend the request. A server that accepts the header without storing anything invites duplicate charges from well-behaved Go clients.
On money paths, servers should stop accepting hope. Require the header on charge, capture, refund, and transfer routes. When it is missing, return 400 Bad Request with a link to the documentation, which is the behaviour the IETF draft describes for an operation that requires a key. Optional headers get skipped by whoever integrates in a hurry, and hurried integrations are the ones that retry most carelessly.
The client's key is half the contract. The server keeps the other half in a key store, and each record holds five things.
| Field | Role |
|---|---|
| Scope and key | Tenant, route, and key string together form the lookup identity |
| Fingerprint | A hash of the request fields that define the intent |
| State | in_flight while the effect runs, sealed once the outcome is known |
| Stored outcome | Status code and response body to replay |
| Expiry | When the record may be deleted |
A record has a short life. It is created in in_flight when a request with an unseen key arrives. It becomes sealed when the outcome is known and stored. It is deleted after expiry. Some requests never begin executing, and their reservations are released rather than sealed, a case the section on replay covers. There are no other transitions. A sealed record never returns to in_flight, and its stored outcome never changes. That immutability is what makes a replay trustworthy.
The fingerprint lets the server notice a key reused with a different request. The IETF draft allows the resource to compute a fingerprint from the payload and lists approaches such as a checksum over the body. Hashing raw bytes is the easy choice and the fragile one. JSON key order, whitespace, and number formatting can change between attempts with no change in intent, especially when a retry passes through a different SDK version or a proxy that re-serialises the body.
A sturdier fingerprint hashes the parsed fields that define the intent, in a fixed order, together with the authenticated principal.
func fingerprint(principal string, r ChargeRequest) string {
canon := fmt.Sprintf("%s|%d|%s|%s", principal, r.Amount, r.Currency, r.Source)
sum := sha256.Sum256([]byte(canon))
return hex.EncodeToString(sum[:])
}Leave out fields that legitimately change between attempts, such as client timestamps, trace identifiers, and retry counters. Include every field whose change would mean a different payment. Including the principal means a leaked key cannot be replayed from another account to read someone else's stored response.
A key string alone is not an identity. Two merchants on one platform can both send order-1001. One client can send the same key to POST /v1/charges and, by mistake, to POST /v1/refunds. If the store looks up by key alone, the first case leaks one tenant's response to another, and the second replays a charge response to a refund request.
Put the tenant, the route, and the key into the unique index. The route component matters most for internal services that call several endpoints with keys derived from the same business identifier. Scope can also be too narrow in the other direction, which the section on failure modes returns to.
On paper the server's steps are simple. Check the store, run the effect, store the outcome. How hard that is depends almost entirely on a question many designs skip. Is the effect a write to the same database as the key store, or a call across a network?
Local effect in one transaction versus remote effect with an in-flight window
If the charge is a row in the same Postgres database that holds the keys, the whole protocol fits in one transaction, and most of the hard cases disappear.
BEGIN;
INSERT INTO idempotency_keys (tenant_id, route, key, fingerprint, state)
VALUES ($1, 'POST /v1/charges', $2, $3, 'sealed');
INSERT INTO ledger_entries (charge_id, tenant_id, amount, currency)
VALUES ($4, $1, $5, $6);
UPDATE idempotency_keys SET status = 201, body = $7
WHERE tenant_id = $1 AND route = 'POST /v1/charges' AND key = $2;
COMMIT;The key row and the ledger row commit together or not at all. A crash before COMMIT leaves nothing behind, so a retry starts clean. A crash after COMMIT leaves a sealed record, so a retry replays. No other transaction ever sees an in-flight state.
Concurrency comes almost free. When two transactions insert the same key, Postgres makes the second insert wait on the unique index until the first transaction ends. If the first commits, the second fails with a unique violation, and the handler reads the sealed record and replays it. If the first rolls back, the second proceeds as though it had arrived alone. That is a wait-until-sealed policy implemented by the database's own locking, with no application code.
Reach for this design whenever you can. Many teams cannot, because the effect that matters, moving money, happens at a payment processor on the far side of an HTTPS call.
A call to a card processor cannot join a local transaction. The protocol splits into three steps, and each boundary between steps is a place where the process can die. The server first reserves the key by inserting it in in_flight and committing. Then it calls the processor. Then it seals the record with the outcome and commits again.
The order is not negotiable. Reserving after the effect opens a gap in which the processor has charged and the store has no record, so a retry looks new and charges again. Reserving first means a crash at any point leaves evidence, a record stuck in in_flight, that a later process can find and finish. Brandur Leach's write-up of Stripe-style keys in Postgres generalises this into atomic phases separated by recovery points, so a request with several external calls resumes from the last completed phase instead of starting over.
The Go below shows the decision logic for the remote case. The store methods wrap the SQL.
func (s *Store) Serve(k Key, fp string, effect func() Outcome) Outcome {
rec, created, err := s.Reserve(k, fp)
switch {
case err != nil:
return Outcome{Status: 503}
case !created && rec.Fingerprint != fp:
return Outcome{Status: 422}
case !created && rec.State == Sealed:
return rec.Outcome.Replayed()
case !created:
return Outcome{Status: 409, RetryAfter: 1}
}
out := effect()
if out.Ambiguous {
return Outcome{Status: 503, RetryAfter: 5}
}
if err := s.Seal(k, out); err != nil {
return Outcome{Status: 503, RetryAfter: 5}
}
return out
}Reserve has to be atomic. It either creates the record or returns the one that exists, and two goroutines must never both see created as true. An ambiguous outcome, such as a timeout from the processor, deliberately leaves the record in flight. Sealing it as a failure would tell the client the payment failed when it may have succeeded. A failed Seal after a successful effect is handled the same way, because the record stays in flight and recovery will finish it.
At this point the server is itself a client, of the processor, and everything this piece says about clients applies to it. If the handler's call to the processor times out and the handler retries, or a recovery job re-sends the call an hour later, the processor faces the same ambiguity the phone did. It needs a key too.
Derive the downstream key deterministically from the upstream one, so every retry and every recovery attempt sends the same value.
func downstreamKey(k Key, step string) string {
parts := strings.Join([]string{k.Tenant, k.Route, k.Value, step}, "\x00")
sum := sha256.Sum256([]byte(parts))
return hex.EncodeToString(sum[:])
}The step component separates calls made within one request, such as an authorisation and a later capture, so they do not collide at the processor. Hashing keeps the client's raw key out of a third party's logs. Never mint a random key per downstream call. That recreates the original bug one hop further in, where it is harder to see.
Chaining keys is what makes the guarantee end to end. The client's key protects the hop from client to API. The derived key protects the hop from API to processor. Any hop without a key is a place where a retry becomes a duplicate.
Retries do not always wait their turn. A mobile stack that times out at five seconds, plus a user who taps twice, can put two requests with one key on the wire within milliseconds. They reach different API instances. If each checks the store, finds nothing, and proceeds, both call the processor.
The reservation must be a single atomic operation against the store. In Postgres it is an insert with a conflict clause.
INSERT INTO idempotency_keys (tenant_id, route, key, fingerprint, state, created_at)
VALUES ($1, $2, $3, $4, 'in_flight', now())
ON CONFLICT (tenant_id, route, key) DO NOTHING
RETURNING key;A returned row means this request owns the key. No row means another request got there first, and the handler reads the existing record to decide what to send back. In Redis the equivalent is SET with the NX and PX options, which sets the value only if the key is absent and attaches an expiry. Redis needs more care on money paths. Its replication is asynchronous, so a failover to a replica that never received the write can lose a reservation, and two requests can each believe they own the key.
A second request that finds a record in flight needs a policy.
| Policy | Client sees | Cost |
|---|---|---|
| Wait for the seal | The original outcome after a delay | Holds a connection and a worker |
| Return 409 at once | A signal to retry with the same key | Clients must loop, and aggressive ones hammer the store |
| Wait briefly, then 409 | Usually the outcome, sometimes 409 | A little of both |
The IETF draft says a server should answer a retry that arrives while the original is still processing with 409 Conflict, and that the client needs to change nothing before retrying. I prefer a short server-side wait of a second or two before that 409. Most in-flight windows are shorter than that, so most duplicates get the real answer on the first try. Unbounded waiting is worse than both options, because one stuck effect then holds every retry's connection open until the load balancer gives up.
A key that arrives with a different fingerprint signals a client bug, and the server must not guess which body the client meant.
Replaying the stored response would tell the client its 8,400 pence payment succeeded when only 4,200 was charged. Applying the new body under the old key would charge twice while the store records one payment. The IETF draft asks for 422 Unprocessable Content here, with a body linking to documentation. Stripe also returns an error when the parameters differ from the original request. Either way, the client must receive an error it cannot mistake for success.
A steady trickle of mismatches in production usually has a small set of causes. A client reuses one key for a whole session. A proxy or SDK mutates the body between attempts. A client changes the amount after a partial failure and retries without minting a new key. Each deserves a ticket to whoever owns that client, and the 422 rate on a dashboard is how you notice them.
Sealing a success is obviously right. The harder question is what to do with every other outcome. The answer turns on whether the effect began and whether the outcome is final.
Decision tree for storing, releasing, or holding an outcome
| Outcome | What the store does | Reason |
|---|---|---|
| Rejected before execution, such as a validation 400 | Stores nothing and releases the reservation | Nothing ran, so a corrected request may proceed |
| Concurrent conflict 409 | Stores nothing for the second request | The first request still owns the key |
| Success 201 | Seals and replays | The standard case |
| Final business failure, such as a declined card | Seals and replays | The same intent should get the same answer |
| Ambiguous, such as a processor timeout | Keeps the record in flight | The money may have moved, and recovery must find out |
Stripe's documented behaviour differs on the last row, and it is worth understanding why. Stripe saves the status code and body of the first request for a key whether it succeeds or fails, and later requests with that key get the same result, including 500 errors. It saves results only once an endpoint has begun executing, so validation failures and concurrent conflicts leave nothing behind and can be retried. That is a coherent contract. A client that receives a replayed 500 learns that retrying with this key will not help, and Stripe runs its own reconciliation behind the error.
For an API you build and operate, I recommend holding ambiguous outcomes in flight instead of sealing them. A sealed 500 is permanent for the life of the key. If the processor did charge, the client can never learn that through this key. It mints a new key, retries, and charges again. Holding the record in flight makes retries return 409 until recovery learns the truth and seals the real outcome, and then the next retry gets the real answer. The cost is that clients see 409 for longer during a processor incident, and you have to build the recovery job described below. I think that cost is worth paying on any route where a false failure leads a user to pay twice.
A declined card deserves a note because it tempts people to release the key. The decline is final for that intent. If the user tries another card, the body changes and so does the intent, and the client should mint a new key. Replaying the decline to a retry of the same body is correct behaviour, not a bug.
Everything so far concerns an HTTP client, but the same contract covers asynchronous work. This is where systems that handle checkout correctly often still produce duplicates.
Brokers deliver at least once by default. A consumer receives a message, performs the effect, and crashes before acknowledging, so the broker redelivers and the consumer performs the effect again. Brokers that advertise exactly-once semantics define the phrase within their own boundary. Kafka's idempotent producer prevents duplicate writes caused by producer retries, and its transactions make a read-process-write cycle atomic when the output goes back into Kafka. Neither reaches a card processor on the far side of an HTTPS call. Amazon SQS FIFO queues drop messages sent with the same deduplication ID within a five-minute interval. That protects the enqueue step for five minutes. It does nothing for the consumer's effect.
So the consumer needs a key, and the source of that key matters. Delivery metadata is a poor source. An SQS receipt handle changes every time the message is received. A message ID survives redelivery of that message, but if the producer's send timed out and the producer sent again, there are now two messages with two IDs for one intent. The sturdier source is a business identifier in the payload that names the intent, such as the order ID plus the action.
A consumer capturing payment for order 7731 should call the charge API with a key derived from order_7731:capture. Redelivery sends the same key. A duplicate publish sends the same key. The charge service collapses both exactly as it collapses a mobile retry.
Resist building a separate deduplication table inside the worker and then calling the charge API without a key. That splits the contract into two halves at two boundaries, and a duplicate slips through whichever boundary the failure happens to cross. One contract belongs at the side-effect boundary. The worker's job is to supply a key that means the same thing every time.
Every key store eventually deletes records, and deletion changes the contract. Stripe states that keys can be removed once they are at least 24 hours old, and that a request reusing a pruned key is treated as a new request. After expiry, a retry is a new payment.
Key expiry compared with the retry windows of different callers
The expiry therefore has to outlast the slowest automated retry of any caller, not the typical one. Retry horizons span several orders of magnitude.
The last point is easy to miss. If recovery re-sends a downstream call with a derived key after the processor has pruned that key, the processor treats the call as new. Recovery must finish well inside the downstream expiry, and alerts should fire long before any in-flight record gets near it.
Storage usually costs less than teams fear. A service handling three million keyed requests a day, at roughly 1.5 KB per record including the response body, holds about 31.5 GB for a seven-day expiry, plus index overhead. That is a modest table for a payments database, and small next to one day of duplicate charges and the support time spent refunding them.
Two rules make expiry safe. First, expire only sealed records. A record still in flight is evidence of an unfinished effect, and deleting it destroys the only pointer to money that may have moved. Second, publish the expiry. The IETF draft requires a resource to publish its idempotency specification, including any expiration policy, and client teams need that number to compare against their own retry horizons.
Processes die between reserve and seal. Deploys kill pods mid-request, nodes fail, and processors time out. Each leaves a record in in_flight with no process working on it. Without cleanup, the client retries until it gives up, and the question of whether money moved stays unanswered.
Recovery sweeper resolving stale in-flight records
A recovery job, running every minute or so, finds in-flight records older than a threshold and resolves each one. Set the threshold well above the longest legitimate effect, for example twice the processor call timeout, so the sweeper never races a live request. Claim each record with SELECT ... FOR UPDATE SKIP LOCKED in Postgres so two sweeper instances never resolve the same one. It is the parallel-retry race again, one level down.
For each stale record, the sweeper asks the downstream system what happened, using the derived key. The simplest way is to re-send the original call with the same derived key. If the processor completed the first call, it replays the stored result and nothing new happens. If the first call never arrived, the processor performs it now. Either way the sweeper gets a final outcome and seals it. Processors that support lookup by key or by your own reference can be queried instead, which avoids performing an effect long after the user left.
That second branch raises a product question the sweeper cannot answer by itself. Should a payment the user abandoned an hour ago still go through? For a checkout, often not, since the user may already have paid another way. For a refund or a payout, almost always yes. Decide per route. When the answer is no, have the sweeper void or reverse the effect rather than simply marking the record failed.
One gauge tells you whether recovery is keeping up, the age of the oldest in-flight record. In a healthy system it stays close to the effect timeout. A rising value means the sweeper is failing or the downstream is down, and it should page someone well before it approaches the downstream key expiry.
A team can adopt the header, build the store, and still ship duplicates. The gaps cluster in a few places, and each one is a gap in mechanism rather than in care.
The key is minted per attempt. An SDK generates a UUID inside its retry loop, so every attempt carries a new key and the server faithfully creates a new charge each time. In the logs, keys never repeat even though retries clearly happen.
The reservation comes after the effect. A handler calls the processor and writes the key afterwards. A crash in between leaves a charge with no record. Code review has to read the order of operations in the handler, not only confirm that a store exists.
The scope is wrong. Scope by key alone and tenants collide. Scope one key to a whole checkout session, and the charge, a tip adjustment, and a partial refund inside that session block or replay each other.
The expiry is shorter than a caller's retry horizon. An offline mobile queue resumes after the key has expired, or a queue redelivers on day four against a 24-hour store.
The downstream hop has no key. The API honours client keys perfectly, then calls the processor with a random key per attempt, so its own retries and its recovery job create duplicates at the processor.
The worker path skips the contract. The checkout path sends keys. The settlement consumer, written by another team a year later, does not. The duplicate appears hours after the purchase, in a batch job nobody associates with checkout.
The store can forget. A cache used as the key store evicts records under memory pressure, or fails over to a replica that missed the reservation. The contract is only as durable as the store behind it.
Fixing any of these means changing the protocol, the scope, the expiry, or the key source. Reminding people to be careful fixes none of them.
The failures above depend on timing, so ordinary tests miss them. A unit test that calls the handler twice in sequence passes against a broken implementation, because the first call finishes before the second starts. The contract needs tests that recreate the timing, and three of them cover most of the risk.
The first test drops the response after the effect commits. Put a fault-injecting proxy such as Toxiproxy between the test client and the API, configured to cut the connection on the response path. The client sees a timeout, retries with the same key, and must receive the original charge ID. Then assert that the ledger holds exactly one row for that key.
The second test sends duplicates in parallel, and it needs nothing more than a shell.
KEY=$(uuidgen)
for i in 1 2 3; do
curl -s -o /dev/null -w "%{http_code}\n" -X POST "$API/v1/charges" \
-H "Idempotency-Key: $KEY" -H "Content-Type: application/json" \
-d '{"amount":4200,"currency":"GBP"}' &
done
waitAcceptable output is one 201 plus replayed 201 responses or 409 responses. Two fresh 201 responses with different charge IDs mean the reservation is not atomic. Run it a few hundred times in CI against a real database, because a race that hides in one run shows up in three hundred.
The third test kills the process between effect and seal. Add a failpoint to test builds, an environment variable that makes the handler exit right after the processor returns. Start a request and let it die. Then confirm that the sweeper finds the in-flight record, resolves it through the derived key, and that a client retry afterwards replays the resolved outcome.
Tests catch regressions before release. Production needs a signal too, and the most useful one is a response header that says a replay happened.
POST /v1/charges HTTP/1.1
Idempotency-Key: checkout_sess_9f2c_pay
Content-Type: application/json
{"amount":4200,"currency":"GBP"}A retry after the first write landed gets the original charge back, marked as a replay.
HTTP/1.1 201 Created
Idempotent-Replayed: true
{"charge_id":"ch_8k1","amount":4200}Stripe marks replayed responses with an Idempotent-Replayed header. The name is not standardised, and the signal matters more than the name. With it, an operator reading logs can tell a prevented duplicate from a request that was only ever sent once. Without it, the store's work is invisible, and a store that quietly stopped working looks exactly like one that never had anything to do.
Four numbers belong on the dashboard for any keyed route. Replay rate shows how often clients retry, and a jump usually points at a client release or a network problem. The 422 rate counts clients misusing keys. The 409 rate shows in-flight contention. The age of the oldest in-flight record shows whether recovery keeps up. Each is a counter or a gauge over the key store, so none of them needs sampling or tracing.
Not every write needs this machinery. Adding it everywhere costs storage, latency on every request, and another system to operate.
Reads belong on GET, which is idempotent by definition. Stripe tells clients not to send keys on GET and DELETE for that reason, since the header has no effect there.
A resource whose identifier the client chooses already carries its own key. PUT /orders/7731 with a full representation is idempotent by method, and repeating it writes the same state. Designing creates this way removes the need for a separate store. In exchange, clients must generate identifiers, and conflicts move onto the identifier.
A two-step create-then-confirm design moves sameness into the state of a resource. Stripe's PaymentIntents follow this shape. The client first creates an intent object representing one payment, then confirms it. The intent's ID becomes the identity of the purchase, and a confirmation against an intent that has already succeeded fails on state rather than charging again. Keys still help on the create call, but a state machine the server owns protects the step that moves money.
Internal tools where a person confirms every action and nothing retries automatically can go without keys, provided a rare duplicate is small and easy to spot. Money paths never qualify. Skipping keys on a charge route because the network seems reliable is a bet against timeouts, and at scale that bet always loses eventually.
Idempotency for an unsafe HTTP method is an agreement about sameness between a caller and a server. It cannot be a server setting, because the server cannot see intent in the bytes it receives. The caller names each intent with a key. The server remembers what it did for that key and gives the same answer every time it is asked.
The idempotency contract as client and server obligations
The client's obligations are few. Mint the key when the user or job commits to the action. Store it with the intent so it survives crashes and restarts. Send it on every retry and never change the body under it. Derive it from a business identifier when the caller is a worker.
The server's obligations are larger. Require the key on money paths. Scope it by tenant and route. Fingerprint the fields that define the intent and reject a mismatch with 422. Reserve atomically before the effect. Keep the whole protocol inside one transaction when the effect is local, and split it into reserve, effect, and seal when it is remote. Give every downstream hop a derived key. Replay final outcomes, hold ambiguous ones in flight, and run a sweeper that resolves them before the downstream forgets its own key. Keep sealed records longer than the slowest caller's retry, publish that expiry, and mark replays so operators can see the store working.
A reviewer testing a claim that some API is idempotent can ask five questions. Which key? Scoped to what? Stored where, and how durably? Kept for how long, compared with the slowest caller? What happens when the process dies between the effect and the seal? A team with crisp answers to all five has built the contract. A team that points to a flag has answered none of them, because a flag cannot see the second post.