From product intent to production evidence
This playbook is a navigation and synthesis layer: use it to choose the next evaluation activity, then follow the linked guides and reusable templates for depth. The LLM Evaluation Framework remains the deep foundational guide; this page focuses on workflow, release decisions, and production practice.
- Define
- Retrieve
- Act
- Validate
- Operate
01 · DefineFoundations
Begin with the product decision and failure cost, not a convenient metric. Translate the intended workflow into observable success criteria, a failure taxonomy, and a versioned golden set. Combine deterministic checks with a review rubric: schemas and exact constraints should be tested by code, while usefulness and ambiguity often need human judgment.
When to use this
At project framing, before replacing a model or prompt, and whenever a new capability changes the system’s risk surface.
Common failure modes
- Testing only happy paths or average quality.
- Letting expected behavior drift after seeing outputs.
- Using vague reviewer criteria without calibration.
Chapters and resources
02 · RetrieveRAG evaluation
Evaluate retrieval and generation as connected but separable layers. Retrieval tests whether required evidence appears and ranks well; context review tests whether supplied passages are relevant and usable; generation tests groundedness, completeness, citations, and justified no-answer behavior. A good final answer cannot diagnose which layer failed.
When to use this
For document question answering, search-grounded assistants, citation systems, or any workflow where answers should be constrained by retrieved evidence.
Common failure modes
- Reporting one blended “RAG score.”
- Checking citations exist but not whether they support claims.
- Omitting answerable and unanswerable cases.
Chapters and resources
03 · ActAgent evaluation
For agents, inspect the process as well as the answer. Define task completion, permitted tools, valid arguments, schemas, action boundaries, and recovery behavior. Trace review reveals silent retries, unnecessary calls, permission mistakes, or a correct final answer produced through an unsafe path. Keep critical constraints as hard gates rather than averaging them into a score.
When to use this
Whenever a system selects tools, changes state, follows a multistep plan, or operates under permissions and structured-output contracts.
Common failure modes
- Scoring only the final response.
- Ignoring malformed or over-broad tool arguments.
- Testing success paths without dependency and permission failures.
Chapters and resources
04 · ValidateReliability and judgment
Treat uncertainty as information about decisions, not decoration. Calibrate confidence-like signals against observed outcomes where possible and evaluate selective answering or escalation. LLM-as-a-judge can scale qualitative review, but its prompt, model, bias, and run-to-run variation become part of the measurement system. Compare automated judgments with blinded human review and track disagreement rather than hiding it.
When to use this
When outputs are subjective, failures have unequal costs, or the system must abstain, escalate, or support reviewer decisions.
Common failure modes
- Treating verbal confidence as calibrated probability.
- Using a judge model as ground truth.
- Resolving reviewer disagreement without examining rubric ambiguity.
Chapters and resources
05 · OperateProduction evaluation
Convert offline evidence into explicit release gates: critical failures, slice thresholds, regression tolerances, reviewer sign-off, monitoring readiness, and rollback criteria. After release, monitor system health and sampled quality, preserve human escalation, and turn new failures into regression cases. Drift analysis should distinguish changes in inputs, retrieval, knowledge, model behavior, tools, and user intent.
When to use this
Before a staged release, after any substantive system change, and continuously for systems whose data, dependencies, or usage can change.
Common failure modes
- Using offline averages as the only release gate.
- Monitoring latency and errors but not output quality.
- Collecting incidents without adding regression tests.
Chapters and resources
A practical starting sequence
- Write five to ten consequential failure modes.
- Create a small golden set with expected behavior and severity.
- Add deterministic validators before model-based grading.
- Choose layer-specific metrics and calibrate reviewers.
- Define release gates, rollback criteria, monitoring, and escalation.
Download evaluation templates Choose RAG metrics
Version and changelog
v1.0 October 1, 2026
- Published the first branded playbook with five-stage evaluation roadmap.
- Added concise guidance for foundations, RAG, agents, reliability, and production.
- Connected the workflow to existing public guides, new templates, and the metric selector.
Future version numbers will change only when substantive content is added.