How do I evaluate an AI agent before production?
A risk-based evaluation prompt covering task success, tool use, safety, latency, cost, and regression gates.
# How do I evaluate an AI agent before production? Act as a AI evaluation lead. ## Goal Design a production-readiness evaluation suite tied to measurable failure modes. ## Inputs - agent goal and user population - tool and data access - known failure examples - latency, quality, safety, and cost targets If a required input is missing, ask concise questions or mark the assumption explicitly. Do not
Use the attached prompt by replacing its input bullets with the real context. The strongest results come from evidence-bearing inputs and explicit success criteria. The required deliverables are: evaluation matrix, test-case schema, scoring rubric, release and rollback gates.