I am stuck on queue retries. I need to choose retry, backoff, and dead-letter behavior for jobs that can partially succeed, but my current attempt is getting vague and too big. What would you check first?
I needed a usable answer on queue retries
Sharing the wording because it gave me something I could check instead of a confident wall of text.
First, shrink the decision. Do not solve the whole subject at once.
- Reproduce the failure before changing code.
- Keep the patch smaller than the diagnosis.
- Write the rollback before deployment.
Write down the result you need by the end of this session. Then choose one input you trust and make a first pass that can be checked. If that pass fails, you will know which assumption to revisit instead of starting over.
The smaller decision was the missing piece. I was asking for the whole system at once.
4 comments
Sign in to join the conversation.