Golden set
Expanded definition
A golden set includes inputs, allowed context, and expected answers or rubrics. It is the opposite of one-off demo prompts. Good sets cover empty retrieval, adversarial documents, tool failures, and policy cases. Store diffs when you change labels. LLM-as-judge can grade at scale but should be calibrated on a human-labeled slice. Public benchmarks are not a golden set for your product.
Related terms
Explore adjacent ideas in the knowledge graph.
Golden set FAQ
What is Golden set?
A golden set is a versioned collection of labeled tasks used to score models and prompts against the behavior you actually want.
How is Golden set used in AI systems?
A golden set includes inputs, allowed context, and expected answers or rubrics. It is the opposite of one-off demo prompts. Good sets cover empty retrieval, adversarial documents, tool failures, and policy cases. Store diffs when you change labels. LLM-as-judge can grade at scale but should be calibrated on a human-labeled slice. Public benchmarks are not a golden set for your product.
Related
Comparisons, tools, and models that connect to this idea.