# Agent release gate checklist

Use this checklist before comparing an agent candidate with the current release.

## Define the contract

- [ ] Give every scenario a stable ID and one task-specific purpose.
- [ ] Declare completion, required tools, forbidden tools, and tool order explicitly.
- [ ] Set step, retry, and supplied-cost limits from the product policy, not this
      fixture.
- [ ] State which citation IDs or other grounding evidence are required.
- [ ] Assign a severity to every invariant and record which severities block release.

## Build comparable evidence

- [ ] Run baseline and candidate against the same policy and scenario inputs.
- [ ] Keep deterministic contract checks separate from model-scored or human review.
- [ ] Preserve every failure reason; do not reduce the result to one average score.
- [ ] Distinguish existing failures from regressions introduced by the candidate.
- [ ] Store the fixture version, policy version, and both release identifiers.

## Protect the evidence

- [ ] Replace user data, credentials, production IDs, and paid requests with synthetic
      values.
- [ ] Save a deterministic manifest and verify that a fresh run byte-matches it.
- [ ] Keep blocking logic in reviewed code rather than a mutable dashboard setting.
- [ ] Record the limits: a deterministic pass does not prove model quality or safety.

## Decide and follow up

- [ ] Block the candidate when a newly introduced failure has a configured blocking
      severity.
- [ ] Review non-blocking failures rather than hiding them from the report.
- [ ] Add confirmed production failures as new cases only after redaction and review.
- [ ] Rerun model-based evaluation and human review where deterministic checks stop.

Licensed under CC BY 4.0. Credit FutureCraft and indicate changes.
