Your application calls another service and the request times out. Did the remote operation fail, or did it finish while the response was lost? This uncertainty is a basic problem in a distributed system: work happens across processes or machines connected by a network.
One component can fail while others remain healthy, and a retry can repeat work that already succeeded. Use deadlines to limit waiting and design repeated operations so duplicates can be recognized or tolerated. These ideas matter for ordinary web applications and queues, not only very large clusters.
Try it in a lab
Model a two-service request with a deadline and bounded retries. Simulate an unavailable dependency and compare immediate failure, retry and queueing behavior using synthetic operations.
Check your understanding
Explain whether an operation can safely be repeated and how duplicate work is detected. Describe what changes when failures are correlated rather than independent.
Before you start
Unlimited retries can amplify an outage. Bound work and preserve enough context to distinguish an uncertain result from a confirmed failure.
Read the official guide
Read Google SRE’s Handling Overload chapter for retry budgets and load amplification. Use the gRPC guides to deadlines and status codes for a concrete example of deadline propagation and why a timeout can follow a successful state change. Record the versions and results of your own exercise. This page proposes a learning activity; it does not report a Kubedex test.
This is a newly written study reference at an address from the original Kubedex course outline. The original lesson was not recovered. It does not include course enrolment, progress tracking or a certificate.
Sources & further reading
Spotted something that needs another look?
Help improve this page →