Reliability and Metastable Failures
If a system is close to overload, with throughput pushed to the limit, it can sometimes enter a vicious cycle in which it becomes less efficient and hence even more overloaded.
For example, if a long queue of requests is waiting to be handled, response times may increase to the point that clients time out and resend their requests. This causes the rate of requests to increase even further, making the problem worse—a Retry Storm. Even when the load is reduced again, such a system may remain in an overloaded state until it is rebooted or otherwise reset. This phenomenon is called a Metastable Failure, and it can cause serious outages in production systems.
To avoid retries overloading a service, you can:
- Exponential Backoff: Increase and randomize the time between successive retries on the client side.
- Circuit Breaker: Temporarily stop sending requests to a service that has returned errors or timed out recently.
- Token Bucket: Use a token bucket algorithm to rate limit requests.
- Load Shedding: Detect when approaching overload and start proactively rejecting requests.
- Backpressure: Send back responses asking clients to slow down.
The choice of queueing and load balancing algorithms can also make a significant difference.
