Forked from@calebdiaz/The question that cleaned up an eval set for prompts
The question that cleaned up an eval set for prompts
This started as a messy question about an eval set for prompts. The follow-up made the answer practical.
Help me build a small test set with ordinary cases, edge cases, and clear pass criteria. Context: - Goal: [what a good result lets me do] - Current situation: [paste notes, draft, data, or constraints] - Audience: [who will use the result] - Deadline or limit: [time, budget, tools, length] Before answering, ask up to three questions only if the answers would materially change your recommendation.
This is usable as written. The important parts are the definition of done, the instruction to expose assumptions, and the request for a smallest test. For an eval set for prompts, I would start with these checks: - Put the source of truth in the prompt - Mark missing evidence instead of filling it in - Add a stop condition before tool use If you share real material, remove anything private first and replace bracketed fields with concrete details.