# Human-in-the-Loop for AI: Approval Gates That Work

> Design human oversight for AI actions with risk tiers, approval state machines, calibrated routing signals, reviewer UX, queue SLAs, audit logs, and release metrics.
> Author: Roman Belov · Published: 2026-04-06 · Source: https://futurecraft.pro/blog/human-in-the-loop/

Human-in-the-loop is often drawn as one box between a model and an action. The box
is the easy part. The difficult part is specifying what the person is responsible
for, what evidence they see, how long they have, and what happens when nobody
answers.

A review button does not transfer accountability to “the human.” The product team
still owns the routing policy, reviewer tooling, audit trail, and failure behavior.

## Start With the Action, Not the Model

The same model output can create very different risk. Drafting a refund reply is
not the same as issuing the refund. Suggesting a database migration is not the
same as running it.

Classify the action on factors the application can observe:

| Factor | Lower concern | Higher concern |
|---|---|---|
| Impact | Formatting, tagging | Money, rights, health, access |
| Reversibility | Easy edit or rollback | Irreversible or costly recovery |
| Scope | One draft | Many users or records |
| Detectability | Error is obvious immediately | Harm appears later or off-platform |
| Authority | User already authorized it | New permission or commitment |
| Time pressure | Can wait for review | Immediate decision required |

The result should be a small set of action policies, not a vague “high-risk” label:

```typescript
interface ActionPolicy {
  actionType: 'draft_reply' | 'send_reply' | 'issue_refund';
  reviewMode: 'none' | 'sample' | 'required';
  maxAmount?: number;
  expiresAfterMs: number;
  timeoutOutcome: 'expire' | 'return_to_user' | 'keep_pending';
  requiredReviewerRole?: string;
  requiredEvidence: string[];
}
```

For regulated or safety-critical use cases, this table is not a substitute for
domain, legal, security, and compliance review. It is the implementation surface
where their requirements become enforceable.

## Use Four Oversight Modes

Most systems need more than “auto” and “manual.”

### 1. Draft for the User

The AI creates an editable artifact but has no permission to execute. The user
remains the actor. This works well for emails, reports, code changes, and form
completion.

The interface must make the boundary honest. A preselected “send immediately”
button can turn nominal review into automation by habit.

### 2. Approval Before Action

The system prepares a typed action and pauses. A qualified reviewer approves,
edits, or rejects it before execution. Use it for irreversible actions, expanded
permissions, significant financial effects, or uncertain policy interpretation.

### 3. Interrupt on Exception

The agent handles a bounded workflow but stops on a policy event: missing required
data, repeated tool failure, conflict between sources, an out-of-distribution
input, or an action outside its delegated scope.

### 4. Post-Action Audit

The action executes, then selected cases are reviewed. This is appropriate only
when errors are detectable and reversible and the residual impact is accepted.
Sampling should include targeted slices and random cases; reviewing only outputs
the model already flagged cannot detect blind spots in the routing signal.

Users also need an **appeal or correction path**. Internal sampling does not help
the person affected by a wrong decision today.

## Confidence Is a Signal, Not Permission

Asking a model to output `confidence: 0.93` does not make the number calibrated.
It may reflect writing style, prompt wording, or a learned preference for high
scores rather than the probability that the action is correct.

Useful routing signals can include:

- deterministic schema and policy checks;
- missing required evidence;
- retrieval coverage and source conflict;
- distance from the evaluated task distribution;
- disagreement between independently designed checks;
- a task score calibrated against labeled outcomes;
- repeated tool or validation failure;
- explicit user request for a person.

Calibrate each signal on representative data. For a score `s`, group examples into
score ranges and compare predicted confidence with observed pass rate. Evaluate by
language, tenant, task subtype, and other slices that change performance. A single
global threshold can hide a weak subgroup.

Threshold selection is a policy decision. Plot, for each candidate threshold:

- volume sent to review;
- harmful error rate among autonomous actions;
- false escalation rate;
- reviewer capacity and wait time;
- outcome by important slice;
- cost of the safe fallback.

Do not optimize “accuracy” when false approval and false escalation have different
consequences.

## Separate Proposal, Approval, and Execution

The model should propose a typed action. It should not embed a side effect inside
the generation call.

```typescript
interface ActionProposal {
  proposalId: string;
  actionType: string;
  arguments: Record<string, unknown>;
  evidenceRefs: string[];
  policyVersion: string;
  createdAt: string;
  expiresAt: string;
  contentHash: string;
}

interface ReviewDecision {
  proposalId: string;
  contentHash: string;
  reviewerId: string;
  decision: 'approved' | 'edited' | 'rejected';
  reasonCode: string;
  decidedAt: string;
}
```

An approval must reference the exact proposal hash. If arguments change, the old
approval is invalid. The executor then validates current permissions, policy,
expiry, and idempotency before acting.

```text
PROPOSED -> PENDING_REVIEW -> APPROVED -> EXECUTING -> EXECUTED
                          \-> REJECTED
                          \-> EXPIRED
```

Make state transitions atomic. A double click, queue redelivery, or retry after a
network timeout must not issue the refund twice. Store one idempotency key per
logical action and record the external system's operation ID.

Approval is not authorization. A reviewer cannot approve an action they would not
be allowed to perform directly.

## Design the Review Queue for Failure

A queue needs more than FIFO and an “urgent” flag.

Each item should define:

- tenant and access scope;
- action type, impact, and affected objects;
- creation and expiry time;
- reviewer role and separation-of-duty requirements;
- evidence references and freshness;
- current workflow state;
- safe timeout outcome;
- deduplication and idempotency keys.

Capacity planning begins with arrival rate, handling time, staffing windows, and
the target wait time. Measure the distribution, not only the average. A queue that
works during office hours may fail every weekend.

When capacity drops, do not silently relax the policy. Safer levers are:

- pause the AI feature or high-risk action;
- return control to the user;
- expire nonessential proposals;
- route to an approved on-call role;
- reduce feature scope;
- process reversible low-risk work under an explicitly reviewed fallback.

The timeout behavior must be visible to users. “Pending review” should not look
like a completed action.

## Build a Reviewer Interface That Supports Judgment

Reviewers need decision-relevant evidence, not a wall of model text.

Show:

- proposed action and affected object;
- before-and-after state;
- authoritative source excerpts with links and timestamps;
- failed rules, missing inputs, and route reason;
- comparable policy examples where appropriate;
- an explicit edit, reject, or request-more-information option;
- the consequences of approval.

Avoid using hidden chain-of-thought as an explanation. A concise evidence summary
and source references are more auditable. Do not let the model's polished rationale
anchor the reviewer before they see the underlying facts.

Reduce fatigue by grouping similar low-risk reviews, rotating assignments, adding
breaks, and monitoring decision time and reversal rate. Seeded quality-control
cases can reveal attention problems, but reviewers should know the program exists
and how results are used.

## Measure the Human and the System

HITL metrics must connect routing, queue health, reviewer decisions, and final
outcomes:

- autonomous, sampled, required-review, expired, and appealed volume;
- harmful error rate by route and action type;
- false escalation and missed escalation;
- queue age percentiles and SLA misses;
- approval, edit, rejection, and reversal rates;
- reviewer agreement on deliberately double-reviewed samples;
- time per decision and changes across a shift;
- incidents caused after approval;
- performance by language, customer segment, and risk slice.

Reviewer disagreement is not automatically a reviewer failure. It may expose an
ambiguous policy or insufficient evidence. Resolve the policy, update the UI, and
relabel affected evaluation cases before tuning the model.

## Turn Corrections Into Controlled Learning

Store structured reason codes and corrected outputs, but do not feed every edit
straight back into a prompt or fine-tuning job. Reviews can be inconsistent,
contain sensitive data, or encode a temporary policy exception.

Use a controlled cycle:

1. redact or restrict sensitive data;
2. adjudicate ambiguous and high-impact examples;
3. version the labeled evaluation set;
4. diagnose the failure source: retrieval, prompt, model, tool, policy, or UI;
5. change one component;
6. run offline evals and shadow tests;
7. canary the new route and retain rollback.

The [agent testing guide](/blog/ai-agent-testing-evaluation/) covers evaluation
sets. If an LLM judge is part of the routing system, calibrate it separately; see
the [LLM-as-judge guide](/blog/llm-as-judge-automated-quality-gate/).

## Audit Without Building a PII Warehouse

An audit record should explain who could do what and why:

- proposal, policy, prompt/workflow, model, and tool versions;
- evidence identifiers and access decisions;
- route reason and validation results;
- reviewer identity and role;
- edits, decision, timestamp, and execution result;
- appeal, reversal, or incident link.

Minimize raw prompts and personal data. Apply retention and deletion policy to
review artifacts. Separate operational access from analytics access, and make
audit logs tamper-evident according to the risk of the system.

## Production Checklist

- [ ] Actions are classified by impact, reversibility, scope, authority, and time.
- [ ] Every action has an explicit oversight mode and safe timeout outcome.
- [ ] Self-reported model confidence is never the sole approval signal.
- [ ] Routing signals are calibrated on representative slices and monitored.
- [ ] Proposal, approval, and execution are separate immutable events.
- [ ] Execution rechecks authorization, expiry, policy, and idempotency.
- [ ] Reviewers see authoritative evidence and can edit, reject, or ask for more.
- [ ] Queue capacity and off-hours behavior are tested before launch.
- [ ] Users can recognize pending review and appeal consequential outcomes.
- [ ] Corrections enter a reviewed eval pipeline, not automatic training.
- [ ] Audit records are useful without retaining unnecessary PII.

## Primary References

- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework)
- [NIST AI RMF Core: human oversight responsibilities](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/)
- [Google PAIR Guidebook: Feedback and Control](https://pair.withgoogle.com/guidebook-v2/chapter/feedback-controls/)
- [OpenAI: A practical guide to building agents](https://openai.com/business/guides-and-resources/a-practical-guide-to-building-ai-agents/)
- [Anthropic: Trustworthy agents in practice](https://www.anthropic.com/research/trustworthy-agents)

Human oversight works when it changes system authority at a precise boundary. If
the person lacks time, evidence, permission, or a safe way to say “no,” the loop
exists only in the architecture diagram.
