GenAIWiki
Evaluation

Faithfulness

Faithfulness is whether every claim in a generated answer is supported by the supplied evidence.

Expanded definition

Faithfulness is an evaluation axis for RAG and extraction: unsupported claims fail even if the prose is fluent. It is distinct from completeness (did we cover the question) and from helpfulness. LLM-as-judge rubrics should score faithfulness separately and list failing quotes. Human spot-checks are still required because judges can miss subtle contradictions.

Related terms

Explore adjacent ideas in the knowledge graph.

Faithfulness FAQ

What is Faithfulness?

Faithfulness is whether every claim in a generated answer is supported by the supplied evidence.

How is Faithfulness used in AI systems?

Faithfulness is an evaluation axis for RAG and extraction: unsupported claims fail even if the prose is fluent. It is distinct from completeness (did we cover the question) and from helpfulness. LLM-as-judge rubrics should score faithfulness separately and list failing quotes. Human spot-checks are still required because judges can miss subtle contradictions.

Related

Comparisons, tools, and models that connect to this idea.