# Deterministic agent evaluation gate: evidence run

Run date: 29 August 2026

## Commands

```bash
deno fmt --check *.ts
deno check agent_gate.ts fixture.ts run_fixture.ts render_manifest.ts \
  agent_gate_test.ts evidence_artifact_test.ts
deno lint *.ts
deno test --allow-read agent_gate_test.ts evidence_artifact_test.ts
deno run run_fixture.ts > agent-evaluation-manifest-2026-08-29.json
deno run --allow-read --allow-write render_manifest.ts \
  agent-evaluation-manifest-2026-08-29.json \
  agent-evaluation-gate-2026-08-29.svg
```

## Result

The dependency-free fixture evaluated six synthetic agent cases. The candidate passed
one case and failed five. It introduced seven invariant failures; six were blocking
because the fixture configuration lists `critical` and `high` as blocking severities.
The medium-severity supplied-cost overrun was recorded but did not block by itself.

The failures were deliberately constructed:

- required account-read tool missing;
- forbidden refund side effect;
- approval requested before the policy lookup;
- required policy citation missing;
- step, retry, and supplied-cost budgets exceeded.

This is a test of the gate, not a benchmark result. The baseline and candidate release
names, policies, tools, costs, and outcomes are synthetic.

## Stored artifact

`agent-evaluation-manifest-2026-08-29.json` byte-matches a fresh deterministic run.
Its SHA-256 at this checkpoint is
`340f801a4f19bd708d224553ab0198903646eb1d82b03ac36e8564385a7963ec`.

`agent-evaluation-gate-2026-08-29.svg` byte-matches a fresh render from the saved
manifest. Its SHA-256 is
`000babe7d7313c01bf1efbce73e47168575c111566d1b7954e13122509e80a44`.
The SVG was rendered at 1280×720 and inspected at native size. All four summary cards,
six case rows, candidate statuses, blocker counts, and evidence limits are visible
without clipping.

Code is licensed under MIT. The checklist, manifest explanation, and diagram are
licensed under CC BY 4.0.

## Limits

The fixture does not call a model, tool, external API, or paid service. It does not
measure model quality, real latency, production cost, judge calibration, or the
statistical power of a dataset. A passing deterministic gate cannot replace
task-specific model evaluation or human review.
