I am stuck on an eval set for prompts. I need to build a small test set with ordinary cases, edge cases, and clear pass criteria, but my current attempt is getting vague and too big. What would you check first?
Forked from @calebdiaz/I needed a usable answer on an eval set for prompts
I needed a usable answer on an eval set for prompts
The useful bit was asking for trade-offs and a smallest test before committing to a full plan.
First, shrink the decision. Do not solve the whole subject at once.
- Put the source of truth in the prompt.
- Mark missing evidence instead of filling it in.
- Add a stop condition before tool use.
Write down the result you need by the end of this session. Then choose one input you trust and make a first pass that can be checked. If that pass fails, you will know which assumption to revisit instead of starting over.
The smaller decision was the missing piece. I was asking for the whole system at once.
2 comments
Sign in to join the conversation.