A client can lose the response to a command even when the server completed the operation. The connection may close after a payment is recorded, a job is created, or an order is accepted. From the client’s point of view, timeout does not reveal whether the command failed before execution or succeeded before the response disappeared.
Retrying is necessary for availability, but repeating a state-changing command can duplicate the side effect. An idempotency key gives all attempts for one logical command the same identity. The server stores the outcome associated with that identity and reuses it when the same command arrives again.
The key does not make arbitrary code idempotent. It creates a deduplication boundary around a command and requires precise rules for key scope, request equivalence, concurrent attempts, retention, and failure recovery.
One logical command keeps one key
The client generates a key before the first attempt and sends the same value on every retry of that command.
POST /payments
Idempotency-Key: 8f54b7d2-5a8e-4f48-9f13-9d44ec6b0d3a
Content-Type: application/json
{
"order_id": "ord_123",
"amount": 4200,
"currency": "USD"
}If the response is lost, the retry carries the same key and the same command payload. A different payment uses a different key, even when every business field happens to match.
This distinction matters. Deduplicating only from order_id, amount, or another business field can merge operations that are legitimately separate. A dedicated key represents retry identity rather than inferred similarity.
The key also needs a scope. A service may define uniqueness per account, API credential, endpoint, or operation type. A compound identity such as (tenant_id, idempotency_key) prevents unrelated tenants from colliding while keeping the lookup local to the relevant security boundary.
The first durable result becomes the replay result
A simple record can hold the command identity and enough information to reproduce its outcome:
idempotency_key
request_fingerprint
status
response_code
response_body
resource_id
created_at
expires_atOn the first request, the server reserves the key and performs the operation. After the operation reaches its durable commit point, the idempotency record stores the completed outcome. A later retry returns that stored result instead of running the side effect again.
first attempt:
key K -> reserve K -> commit operation -> store result R
retry:
key K -> find completed K -> return RThe replay response does not have to be byte-for-byte identical if the API contract permits regenerated headers or transport metadata. The important property is that the logical operation is not applied twice and the returned status remains consistent with the original committed outcome.
Storing only a resource identifier can be enough when the response can be reconstructed safely. Storing the original response is useful when reconstruction could observe newer state and produce a materially different representation.
Request equivalence must be checked
A key should not silently authorize a different command.
Suppose a client first sends:
{"order_id":"ord_123","amount":4200,"currency":"USD"}and later reuses the same key with:
{"order_id":"ord_123","amount":8400,"currency":"USD"}Returning the first result hides a client defect. Executing the second request defeats deduplication. A safer contract rejects reuse of the key with different operation parameters.
Services often store a fingerprint derived from the canonical request fields that define the operation. Canonicalization must be deliberate. Hashing raw JSON bytes can treat harmless formatting or object-member order as different requests, while omitting a meaningful field can treat distinct commands as equivalent.
Authentication context may also belong outside or inside the fingerprint depending on the key scope. The rule must match the API’s authorization and tenancy model rather than relying on a generic hash recipe.
Concurrent duplicates need one winner
Retries are not always sequential. A client, proxy, or job runner can issue overlapping attempts before the first one finishes.
A read-then-insert sequence is unsafe:
A: lookup K -> absent
B: lookup K -> absent
A: perform side effect
B: perform side effectThe reservation for a key needs an atomic conflict boundary, commonly a database unique constraint or another compare-and-create primitive.
INSERT INTO idempotency_keys (tenant_id, key, status)
VALUES ('tenant_7', '8f54b7d2...', 'in_progress')
ON CONFLICT (tenant_id, key) DO NOTHING;Only the request that creates the reservation proceeds as the owner. A concurrent request that finds in_progress needs a documented response policy: it can wait for completion, poll internal state, or return a retryable status. Running the command in parallel is not a valid fallback if duplicate effects are forbidden.
The reservation mechanism and the side effect must still share a reliable commit strategy. Winning a unique insert alone does not solve a crash between reserving the key and changing business state.
The atomic boundary decides crash behavior
The strongest design places the idempotency record and business mutation in the same database transaction when both live in one transactional store.
BEGIN;
INSERT INTO idempotency_keys (...);
INSERT INTO payments (...);
UPDATE orders
SET payment_status = 'paid'
WHERE id = 'ord_123';
UPDATE idempotency_keys
SET status = 'completed',
response_code = 201,
resource_id = 'pay_456'
WHERE key = '8f54b7d2...';
COMMIT;A rollback removes both the reservation and the business mutation. A commit makes both visible together.
When the side effect crosses a transactional boundary, the design needs another recovery mechanism. Calling an external provider after reserving a local key can leave an in_progress row if the process stops. The service then needs enough durable state to reconcile that attempt without blindly repeating an uncertain external effect.
This is the same distributed-systems constraint seen in other dual-write problems: a deduplication table cannot create atomicity across independent systems. The external operation may need its own idempotency token, a durable workflow, an outbox-style handoff, or reconciliation based on a stable operation identifier.
Failure responses need an explicit policy
Not every response should be cached under the key.
A validation error detected before any side effect may be safe to return again, but some APIs permit the caller to correct the request and retry with a new key. A server failure before the command starts may leave the key unused. A failure after commit must preserve the committed outcome even if response serialization or network delivery fails.
The useful boundary is not simply HTTP 2xx versus 5xx. It is whether the logical command crossed a durable point that must not be repeated.
An implementation can model states such as:
reserved -> completed
reserved -> failed_final
reserved -> recoverableState names are less important than deterministic transitions and recovery rules. An ambiguous state that causes operators or workers to rerun side effects manually can bypass the protection the key was intended to provide.
Retention defines the deduplication window
Idempotency records cannot usually remain forever. Their retention period sets the interval during which a repeated key is guaranteed to resolve to the earlier operation.
Deleting records too soon can turn a delayed retry into a new command. Keeping them indefinitely increases storage, index size, and data-retention obligations.
The API contract should state the supported key lifetime when clients can act on it. Cleanup should use the same time semantics as the lookup path, and expiration should not race with an active operation. A record marked in_progress normally needs separate handling from an old completed record.
Client retry policy and server retention belong together. A client that may retry for 24 hours cannot rely on a server that forgets keys after 30 minutes.
Keys are not authorization credentials
An idempotency key identifies a command attempt; it should not grant access to the command or its result.
The server must perform normal authentication and authorization checks on retries. Key lookup also needs tenant isolation so one caller cannot probe another caller’s stored outcomes by guessing or obtaining a key.
High-entropy keys reduce accidental collision and guessing risk, but entropy does not replace access control. Logs and metrics should avoid exposing sensitive request bodies merely because they are attached to deduplication records.
Rate limits also remain relevant. A client can send the same key repeatedly and consume lookup, waiting, or response bandwidth even though the business mutation runs once.
Observability should expose deduplication states
Useful metrics separate first attempts from replays and conflicts. A service can track new reservations, completed replays, payload mismatches, concurrent in_progress hits, expired-key reuse, recovery actions, and age of the oldest unfinished reservation.
These signals reveal different faults. A spike in replays can indicate network instability or aggressive client retries. Many payload mismatches can indicate a client that reuses keys incorrectly. Old unfinished reservations point to crash recovery or workflow stalls.
Tracing should keep the logical operation identifier stable across retries while preserving a distinct request or span identity for each transport attempt. That separation makes one business command visible without collapsing several network attempts into one event.
Idempotency is a contract around retries
An idempotency key turns an uncertain retry from “run this command again” into “resolve the outcome of this command identity.” That shift is small at the HTTP surface and substantial in storage semantics.
Reliable implementations define the key scope, compare request meaning, serialize concurrent attempts, align the key record with the business commit, retain records for a stated interval, and recover ambiguous external effects without speculative repetition.
The mechanism is most valuable where retries are expected and duplicate side effects are expensive. It does not remove distributed failure modes; it gives repeated attempts a durable identity that the service can use to contain them.