Deterministic decision aid

RAG Evaluation Metric Selector

Answer seven focused questions to assemble a metric stack for your RAG system. The selector runs locally, calls no model or API, and does not store your answers.

Your evaluation context

1. What are you evaluating?
2–6. What evidence and risks apply?
Suggested minimum stack

Recommendations

This is a starting point. Set thresholds by risk, validate automated graders against human review, and inspect results by important segments.

Evaluation layers

Retrieval and generation answer different questions

Retrieval evaluation

Retrieval metrics test whether the system found useful evidence and ranked it well. Recall@k, Precision@k, MRR, and nDCG require relevance labels. They do not measure whether the generator used the evidence faithfully or produced a correct answer.

Generation evaluation

Generation metrics examine the answer: correctness when references exist, groundedness in supplied context, completeness, and appropriate refusal. A grounded answer can still rely on an incorrect or stale source, so source quality remains a separate control.

Evidence design

Reference-based and reference-free checks

When references exist

Reference answers or required facts support answer-correctness checks. Relevance labels support retrieval recall and ranking metrics. References should be reviewed, versioned, and broad enough to allow valid alternative wording.

Without references

Context relevance, groundedness, completeness, and human review can still reveal important failures. These checks do not establish correctness on their own: a response may be faithful to incomplete evidence or receive an inconsistent automated judgment.

How to interpret your generated stack

The minimum stack highlights the smallest cross-layer set suggested by your answers. The detailed cards add diagnostic metrics and state what each one does not measure. Use the result to design a versioned test set—not as a universal score or automatic release decision.

Use metrics as a diagnostic system

No single score is universally sufficient. Retrieval metrics cannot prove generation quality; groundedness cannot prove source truth; end-to-end correctness can hide the failing layer. Pair metrics with a versioned golden set, failure taxonomy, slice analysis, and explicit release gates.