Back to glossary
AI quality and safety

Evaluation set

A versioned collection of representative questions and expected behavior used to test retrieval and generated answers consistently.

Business example

An HR team tests 30 approved questions against a fixed policy pack. Each row records the expected source, required facts, access role, and whether the assistant should answer, clarify, or decline.

How it works

  1. 1.Select representative questions from the intended users and workflows.
  2. 2.Map each question to expected sources, facts, access rules, and an expected outcome.
  3. 3.Include difficult cases such as paraphrases, obsolete sources, permission denials, conflicts, and missing knowledge.
  4. 4.Run the same version against the system and score retrieval separately from the generated answer.
  5. 5.Review failures with domain owners, revise sources or the system, and preserve dataset versions for comparison.

Common misconceptions

A golden dataset is a permanent collection of perfect model answers.
It is a reviewed reference for a defined source and use-case version; it must change when policies, workflows, or risks change.
A large test set is automatically representative.
Many duplicate easy questions can miss permission, freshness, ambiguity, and no-answer cases that determine business safety.

Related concepts

Related reading

Sources

Reviewed:

Reviewed by: Javier Chulvi Bernad · LLM Engineer · Madrid