Reproducible agent-evaluation evidence · MIT + CC BY 4.0
Deterministic agent release gate
Compare a candidate agent with its baseline case by case. The fixture checks completion, tool policy, order, grounding, steps, retries, and supplied cost without calling a model or an external API.
Files
Per-case contracts
Each scenario owns its required tools, forbidden effects, order, grounding, and execution budgets. There is no universal score threshold.
Regression-aware gate
Existing failures remain visible, while only new failures with configured critical or high severity block this candidate.
Inspectable output
The manifest keeps every reason in stable order, and both the JSON and SVG byte-match a fresh deterministic run.
Verified scope
Nineteen tests cover each failure code, invalid policies, duplicate cases, stable ordering, baseline comparison, saved JSON, saved SVG, licensing, and credential or direct-PII patterns. The fixture uses six synthetic cases and no paid request.