Reproducible agent-evaluation evidence · MIT + CC BY 4.0

Deterministic agent release gate

Compare a candidate agent with its baseline case by case. The fixture checks completion, tool policy, order, grounding, steps, retries, and supplied cost without calling a model or an external API.

This deliberately failing candidate proves the gate mechanics. It does not measure model quality, real latency, production cost, dataset coverage, or judge calibration.
Six synthetic agent cases with one pass, seven new failures, and a blocked release gate

Files

Per-case contracts

Each scenario owns its required tools, forbidden effects, order, grounding, and execution budgets. There is no universal score threshold.

Regression-aware gate

Existing failures remain visible, while only new failures with configured critical or high severity block this candidate.

Inspectable output

The manifest keeps every reason in stable order, and both the JSON and SVG byte-match a fresh deterministic run.

Verified scope

Nineteen tests cover each failure code, invalid policies, duplicate cases, stable ordering, baseline comparison, saved JSON, saved SVG, licensing, and credential or direct-PII patterns. The fixture uses six synthetic cases and no paid request.