Suggested minimum stackRecommendations
This is a starting point. Set thresholds by risk, validate automated graders against human review, and inspect results by important segments.
Recommended metrics and their limits
Evaluation layersRetrieval and generation answer different questions
Retrieval evaluation
Retrieval metrics test whether the system found useful evidence and ranked it well. Recall@k, Precision@k, MRR, and nDCG require relevance labels. They do not measure whether the generator used the evidence faithfully or produced a correct answer.
Generation evaluation
Generation metrics examine the answer: correctness when references exist, groundedness in supplied context, completeness, and appropriate refusal. A grounded answer can still rely on an incorrect or stale source, so source quality remains a separate control.
Evidence designReference-based and reference-free checks
When references exist
Reference answers or required facts support answer-correctness checks. Relevance labels support retrieval recall and ranking metrics. References should be reviewed, versioned, and broad enough to allow valid alternative wording.
Without references
Context relevance, groundedness, completeness, and human review can still reveal important failures. These checks do not establish correctness on their own: a response may be faithful to incomplete evidence or receive an inconsistent automated judgment.
How to interpret your generated stack
The minimum stack highlights the smallest cross-layer set suggested by your answers. The detailed cards add diagnostic metrics and state what each one does not measure. Use the result to design a versioned test set—not as a universal score or automatic release decision.
- Set thresholds by task risk and inspect important segments separately.
- Validate model-based graders against human review; an LLM judge is not ground truth.
- Keep citation coverage separate from citation correctness: citation presence does not prove support.
- In production, combine sampled quality review with no-answer, unsupported-answer, escalation, latency, and error rates.
Use metrics as a diagnostic system
No single score is universally sufficient. Retrieval metrics cannot prove generation quality; groundedness cannot prove source truth; end-to-end correctness can hide the failing layer. Pair metrics with a versioned golden set, failure taxonomy, slice analysis, and explicit release gates.