Evaluation set
A versioned collection of representative questions and expected behavior used to test retrieval and generated answers consistently.
Business example
An HR team tests 30 approved questions against a fixed policy pack. Each row records the expected source, required facts, access role, and whether the assistant should answer, clarify, or decline.
How it works
- 1.Select representative questions from the intended users and workflows.
- 2.Map each question to expected sources, facts, access rules, and an expected outcome.
- 3.Include difficult cases such as paraphrases, obsolete sources, permission denials, conflicts, and missing knowledge.
- 4.Run the same version against the system and score retrieval separately from the generated answer.
- 5.Review failures with domain owners, revise sources or the system, and preserve dataset versions for comparison.
Common misconceptions
Related concepts
Related reading
Sources
- https://learn.microsoft.com/en-us/azure/databricks/agents/tutorials/ai-cookbook/fundamentals-evaluation-monitoring-rag
- https://cloud.google.com/blog/products/ai-machine-learning/optimizing-rag-retrieval?hl=en
- https://docs.aws.amazon.com/bedrock/latest/userguide/evaluation.html
Reviewed:
Reviewed by: Javier Chulvi Bernad · LLM Engineer · Madrid