Open templates

Applied AI Resources & Template Library

Download practical starting points for evaluation and experimentation. No email gate, no account, and no proprietary data or scoring rules.

These v1.0 resources are platform-neutral and intentionally lightweight. Every example row is synthetic and generic. Adapt the fields, thresholds, reviewers, and release gates to the risk of your own approved system.

CSV · synthetic examples

Golden Test Set Template

Seed a versioned regression set for factual answers, no-answer behavior, structured output, tool selection, and safety boundaries.

Intended use: teams defining test inputs and explicit expected behavior before release.

test_id, category, input, expected_behavior,
must_include, must_not_include, severity, review_notes
Download CSV

Related: golden test set guide

YAML · reusable rubric

Agent Evaluation Rubric

Review task completion, tool use, constraints, schema validity, groundedness, recovery, escalation, and efficiency.

Intended use: structured human review and reviewer calibration.

dimensions:
  - task_completion
  - tool_selection
  - schema_validity
  - groundedness
  - safety_escalation
  - review_outcome
Download YAML

Related: agent evaluation beyond prompting

Markdown · release workflow

AI Agent Release Test Plan

A platform-neutral plan covering environments, golden sets, negative tests, failure handling, review, gates, rollback, and monitoring.

Intended use: planning pre-release validation for tool-using agents.

## Golden-set testing
## Negative tests
## Tool and API failures
## Permission failures
## Release gates
## Rollback criteria
## Monitoring
Download Markdown

Related: testing and evaluating agents

Markdown · experiment planning

Experiment Design Checklist

Define hypothesis, metrics, randomization, sample size, MDE, power, duration, stopping rules, data quality, and the decision rule.

Intended use: experiment review before exposure begins.

[ ] Hypothesis and decision
[ ] Primary metric and guardrails
[ ] Randomization unit
[ ] Sample size, MDE, and power
[ ] Stopping and decision rules
[ ] Data-quality checks
Download Markdown

Related: A/B sample-size calculator

Put the templates into a workflow

Start with the Applied AI Evaluation Playbook for a production-oriented sequence, use the RAG Metric Selector to choose metric families, and browse all interactive tools. If these resources are useful, the newsletter is optional and always comes after the download.