These v1.0 resources are platform-neutral and intentionally lightweight. Every example row is synthetic and generic. Adapt the fields, thresholds, reviewers, and release gates to the risk of your own approved system.
CSV · synthetic examplesGolden Test Set Template
Seed a versioned regression set for factual answers, no-answer behavior, structured output, tool selection, and safety boundaries.
Intended use: teams defining test inputs and explicit expected behavior before release.
test_id, category, input, expected_behavior,
must_include, must_not_include, severity, review_notes
Download CSVRelated: golden test set guide
YAML · reusable rubricAgent Evaluation Rubric
Review task completion, tool use, constraints, schema validity, groundedness, recovery, escalation, and efficiency.
Intended use: structured human review and reviewer calibration.
dimensions:
- task_completion
- tool_selection
- schema_validity
- groundedness
- safety_escalation
- review_outcome
Download YAMLRelated: agent evaluation beyond prompting
Markdown · release workflowAI Agent Release Test Plan
A platform-neutral plan covering environments, golden sets, negative tests, failure handling, review, gates, rollback, and monitoring.
Intended use: planning pre-release validation for tool-using agents.
## Golden-set testing
## Negative tests
## Tool and API failures
## Permission failures
## Release gates
## Rollback criteria
## Monitoring
Download MarkdownRelated: testing and evaluating agents
Markdown · experiment planningExperiment Design Checklist
Define hypothesis, metrics, randomization, sample size, MDE, power, duration, stopping rules, data quality, and the decision rule.
Intended use: experiment review before exposure begins.
[ ] Hypothesis and decision
[ ] Primary metric and guardrails
[ ] Randomization unit
[ ] Sample size, MDE, and power
[ ] Stopping and decision rules
[ ] Data-quality checks
Download MarkdownRelated: A/B sample-size calculator