Hedged Requests Trade Duplicate Work for Lower Tail Latency
Most requests may finish quickly while a small fraction take much longer. A busy worker, a transient network queue, a cold cache entry, garbage collection, storage contention, or another local disturbance can stretch one attempt far beyond the median. At scale, those slow outliers become visible in p95, p99, and higher-percentile latency even when average service time looks healthy.
A hedged request sends an additional attempt after the original has been outstanding for a chosen delay. Both attempts represent the same logical operation. The caller accepts the first valid result and cancels or ignores the remaining attempt.
The mechanism spends extra work to reduce exposure to an unusually slow execution path. That trade can be useful for idempotent reads with enough spare capacity. Applied indiscriminately, it can amplify load at exactly the moment a service is already struggling.
A hedge is delayed, not immediate duplication
Sending two copies of every request at the same time can reduce latency, but it also approaches a 2x request multiplier. Hedging usually waits before creating the second attempt.
A simple policy looks like this:
primary = send(request)
wait until:
primary completes
or hedge_delay expires
if primary completed:
return primary.result
hedge = send(request)
result = first_success(primary, hedge)
cancel_other_attempt()
return resultThe delay is central to the design. Requests that complete within the normal range use one attempt. Only requests that cross the delay consume a hedge.
Suppose the hedge delay is near a recent p95 for the operation. Roughly the slowest slice of requests becomes eligible for a second attempt, rather than the entire workload. The exact fraction depends on the live latency distribution, timer precision, retries, queueing, and how the threshold is maintained.
This is different from a timeout. A timeout places an upper bound on waiting and may terminate an attempt. A hedge starts another attempt while the first can still succeed. The two mechanisms can coexist, but they control different parts of the request lifecycle.
Tail latency depends on correlated failure modes
A hedge helps when the original and duplicate are unlikely to encounter the same source of delay. If one worker is paused while another is healthy, routing the hedge to a different worker can avoid the pause. If every replica is waiting on the same saturated database, the duplicate is likely to encounter the same bottleneck.
Correlation therefore matters more than the number of attempts alone.
A useful hedge path tries to diversify conditions that can produce an outlier:
- a different replica when routing permits it;
- a separate connection when one connection can suffer head-of-line blocking;
- an independently scheduled worker;
- a replica with equivalent data and consistency semantics.
Diversification must not silently change correctness. Reading from a different replica is only safe when that replica satisfies the operation’s consistency requirements. A faster stale response is still the wrong response when fresh state is required.
Hedging also cannot repair deterministic expensive work. If every attempt must scan the same large dataset or wait for the same serialized critical section, duplication adds cost without creating an independent path to completion.
The first result is not always the first acceptable result
“First wins” is convenient shorthand, but production logic usually needs a more precise rule.
If the primary returns a fast transport error while the hedge is still running, immediately returning that error may discard a successful response that is milliseconds away. A caller can instead prefer the first successful result while retaining defined rules for terminal errors.
The policy should specify:
success + pending -> return success, cancel pending
success + error -> return success
terminal error + pending -> wait according to policy
two terminal errors -> return selected error
deadline reached -> stop all attemptsError classification is part of the contract. Authentication failures, malformed requests, or application-level validation errors generally do not become healthier through duplication. Transient transport failures may justify waiting for another already-running attempt.
The caller also needs one overall deadline. Giving each hedge a fresh full timeout can extend total latency beyond the budget of the logical request. Every attempt should inherit the remaining time from the same request deadline.
Cancellation saves capacity only when it propagates
Once one attempt succeeds, the other attempt has no value to the caller. Cancelling it promptly can reduce wasted CPU, network traffic, database work, and connection occupancy.
Cancellation is not automatically equivalent to stopping work. The client library may cancel a future while bytes are already in flight. The server may have started a query that does not observe client disconnects. A downstream database may continue executing after the application abandons the result.
The effective cost of hedging depends on how far cancellation propagates through the stack.
Instrumentation should distinguish at least:
- hedges issued;
- hedges that produced the winning result;
- losing attempts cancelled before dispatch;
- losing attempts cancelled during execution;
- losing attempts that ran to completion;
- added downstream requests and service time.
A high hedge-win rate can indicate useful protection from outliers, but it can also indicate that the primary path is routinely slow. Hedging should not conceal a persistent capacity or routing defect.
Duplicate writes need stronger semantics
Read-only operations are the simplest candidates because duplicated execution usually has no externally visible side effect. Writes are different.
If two attempts can both create a charge, append an event, send a message, or mutate state, cancellation is insufficient. The losing request may already have committed before the caller sees the winner.
A write hedge requires an operation-level idempotency mechanism or another deduplication protocol with semantics that cover concurrent attempts. A stable idempotency key can let the receiving system recognize both attempts as one logical operation, but only if the key is enforced at the boundary where the side effect becomes durable.
Even then, hedging a write deserves a specific justification. Duplicate execution can increase lock contention and complicate observability. The latency benefit must be weighed against the mutation path’s correctness and capacity constraints.
Capacity limits belong beside the hedge policy
Hedging consumes reserve capacity. During normal operation that reserve may be cheap. During overload, duplicate work can create a positive feedback loop:
service slows
-> more requests cross hedge delay
-> more duplicate work arrives
-> queues grow
-> service slows furtherA safe implementation bounds the extra traffic independently from ordinary traffic. Common controls include a hedge budget, a concurrency limit for hedged attempts, and suppression when the downstream reports saturation.
For example, a caller can allow hedges only while both conditions hold:
hedges_in_flight < hedge_concurrency_limit
hedges_issued / original_requests < hedge_rate_budgetThe budget should be measured over a defined window rather than treated as an abstract percentage. A burst of hedges can matter even when the long-term average looks small.
Admission control also needs to decide which work gets priority. In many systems, an original request should have access to capacity before a speculative duplicate. Separate pools or lower scheduling priority for hedges can preserve that distinction.
The delay should follow the operation, not the whole service
One global hedge delay is rarely appropriate. Different endpoints and request classes can have very different service-time distributions.
A 40 ms delay might be late for a cache lookup and aggressive for a complex search. Mixing them into one percentile produces a threshold that represents neither operation well.
The delay can be derived from recent latency for a stable request class, with safeguards against abrupt movement and sparse samples. It can also be configured from an explicit service-level budget. In either case, the metric should represent the same stage that the hedge policy controls.
Queue time deserves special attention. If the client measures end-to-end latency but the server metric begins only after a worker starts processing, the chosen threshold can miss queueing that dominates user-visible delay.
A hedge delay is therefore an operational parameter, not a universal constant. It should be observable, bounded, and changed with the same care as other load-control settings.
Measure the latency gain against the work multiplier
A successful rollout needs both sides of the trade.
Latency metrics can compare p50, p95, p99, and deadline-exceeded rates before and after hedging. Capacity metrics should track the additional request rate, CPU time, database queries, bytes transferred, connection occupancy, and cancellation effectiveness attributable to hedges.
A useful ratio is:
work multiplier =
total physical attempts
/ logical requestsA value near 1 means hedging is rare. A rising multiplier means more logical requests are consuming multiple executions. The acceptable value depends on spare capacity and the cost of each attempt; there is no safe universal target.
Testing should include overload, not only healthy steady state. A policy that improves p99 under light load can destabilize a saturated dependency if its hedge rate rises as latency rises.
Hedged requests are most effective when slow attempts are occasional, execution paths have enough independence, operations are safe to duplicate, cancellation is meaningful, and spare capacity is explicitly budgeted. Under those conditions, speculative work can convert a small amount of redundant execution into a narrower tail without turning every request into permanent duplication.