I needed a usable answer on a flaky test nobody can reproduce
This started as a messy question about a flaky test nobody can reproduce. The follow-up made the answer practical.
I am stuck on a flaky test nobody can reproduce. I need to isolate a timing-dependent test failure and design a deterministic reproduction, but my current attempt is getting vague and too big. What would you check first?
First, shrink the decision. Do not solve the whole subject at once. 1. Reproduce the failure before changing code. 2. Keep the patch smaller than the diagnosis. 3. Write the rollback before deployment. Write down the result you need by the end of this session. Then choose one input you trust and make a first pass that can be checked. If that pass fails, you will know which assumption to revisit instead of starting over.