> The instinct during an availability incident is to add headroom: raise CPU limits, increase the worker pool, add replicas. That can help with genuine capacity problems. Here, it would only give the retry loop more workers to occupy.
This is something during production issues I have a really difficult time sometimes communicating to peers. "Our services are timing out, we're seeing high latency, increase all resources!" is the knee jerk response, but sometimes, and even often, if the underlying cause of the degradation is something like, a database locking up, increasing workers and giving them more firepower might just make the situation even worse. It happens a lot more than you would think.
Your comment reminded me of rachaelbythebay.com. While her site seems to be down at the moment, she has shared some useful war stories from debugging similar scenarios. I understand systems better after having read her work.