Insight AI & Agents

Human-in-the-Loop AI: Designing Approval Gates That Actually Control Risk

An engineering guide to risk tiers, durable approval state, reviewer UX, audit trails, and safe execution for consequential AI agent actions.

A hand authorizes one sealed cartridge on an execution rail while another remains outside the gate.

Reviewed: August 2026.

A human approval button is useful only when it sits on the execution path. If the agent has already sent the payment, deleted the record, or changed production before asking, the button is theater. A defensible human-in-the-loop architecture separates planning from execution, persists the exact proposed action, checks policy, presents the consequential details to an authorized reviewer, and executes only the approved version.

This is not an argument for putting a person in front of every AI response. Routine, reversible, low-impact work can often proceed automatically. The goal is to spend human attention where uncertainty and consequence intersect: financial actions, destructive operations, access changes, production deployments, external communications, and exceptions to established policy.

Approval is an authorization control, not a confidence score

Models can estimate, classify, and recommend, but a model’s confidence is not authority. A request may be understood perfectly and still be prohibited. Conversely, a low-confidence formatting task may be harmless. Approval policy should therefore be based on the action, resource, user, value, reversibility, and organizational rules rather than a single model-generated confidence number.

Amazon Bedrock Agents, for example, can require user confirmation before an action-group function is invoked. AWS describes this as a safeguard against unintended actions, including those induced by prompt injection. That is a useful platform mechanism, but the application still owns identity, business policy, reviewer authorization, audit retention, and the service-side checks that occur when execution begins.

NIST’s AI Risk Management Framework likewise treats human-AI roles, responsibilities, and oversight as governance concerns. The implication for engineering teams is concrete: name who may propose, approve, execute, cancel, investigate, and change the policy. “A human reviews it” is not an operating model.

Use risk tiers to preserve attention

An approval system that interrupts people for everything will be bypassed or rubber-stamped. Start by classifying capabilities, not individual prompts. Each tool operation receives a default tier, and deterministic policy can raise the tier based on amount, environment, sensitivity, scope, or unusual context.

Tier Typical examples Default handling
Tier 0: Informational Search approved documents, summarize a supplied file, read public status Automatic, with normal logging and access controls
Tier 1: Low-impact reversible Draft an internal note, add a label, create a nonpublic ticket Automatic when policy passes; notify or sample for review
Tier 2: Material or externally visible Send a customer message, modify a record, change a schedule Preview and explicit approval by an authorized user
Tier 3: High-impact or difficult to reverse Move money, delete data, change permissions, deploy to production Strong approval, separation of duties where required, and enhanced audit
Prohibited Actions outside policy, unsupported legal authority, or unbounded destructive access Do not expose the capability to the agent

The tiers should be specific to the organization. A $25 credit may be routine for one business and material for another. A staging deployment may be Tier 1 while production is Tier 3. Deleting a generated draft differs from deleting a customer account. The policy engine needs those distinctions in data, not buried in a prompt.

The approval architecture

A robust flow has two trust boundaries. The first separates model planning from the application. The second separates an approved request from the privileged executor.

  1. Interpret: the model proposes a named operation and structured arguments. It has no direct credential to the target system.
  2. Validate: application code checks schema, resource identifiers, allowed ranges, user intent, tenant boundaries, and prerequisites.
  3. Classify: deterministic policy assigns a risk tier and identifies required approver roles.
  4. Persist: the system stores an immutable proposal snapshot, a digest of consequential parameters, provenance, expiration, and status.
  5. Present: the reviewer sees what will happen, where, to whom, with what value, and whether it is reversible.
  6. Authorize: the approval service verifies reviewer identity, role, separation-of-duties rules, freshness, and any step-up authentication.
  7. Execute: a narrow service revalidates the approved snapshot and invokes the target API with a short-lived workload identity.
  8. Record: the audit event links proposal, approval or rejection, execution result, model and policy versions, and correlation identifiers.

The agent should not be able to edit the proposal after approval. If a material field changes, create a new proposal and require a new decision.

What an approval record must contain

The record below is illustrative application data, not a vendor API. It shows the difference between a chat message saying “approved” and an auditable authorization decision.

{
  "proposal_id": "apr_01JEXAMPLE9K2Q",
  "status": "pending",
  "requested_by": "user:8f3a",
  "operation": "production.change.apply",
  "resource": "service:orders-api",
  "environment": "production",
  "parameters": {
    "release_id": "release-2026-08-27.3",
    "change_ticket": "CHG-1842"
  },
  "parameter_digest": "sha256:REPLACE_WITH_COMPUTED_DIGEST",
  "risk_tier": 3,
  "required_approver_roles": ["production-change-approver"],
  "created_at": "2026-08-27T14:20:00Z",
  "expires_at": "2026-08-27T14:35:00Z",
  "policy_version": "prod-change-policy-12",
  "correlation_id": "req_6c924e"
}

Do not store secrets in this record or in the approval UI. Reference a managed secret by identifier if the executor needs one. The digest helps the executor prove that the approved parameters are the parameters it received, but it does not replace access control or a canonical serialization scheme. Use a tested library and a defined encoding when signatures or digests carry security meaning.

Design the button around a safe state machine

Approval is asynchronous. The requester may close the browser, a reviewer may respond minutes later, or the target system may change while the proposal waits. Persist state outside the model session and make transitions explicit:

draft → validated → pending → approved | rejected | expired | canceled → executing → succeeded | failed | reconciliation_required

Allow only valid transitions. Approval from pending is possible; approval from expired is not. Execution should claim an approved proposal atomically so two workers cannot perform it twice. Use an idempotency key at the target API whenever available. If the target returns an ambiguous timeout, do not blindly repeat a financial or destructive operation. Reconcile by the idempotency key or target transaction identifier.

AWS Step Functions demonstrates the general callback pattern: a workflow pauses, waits for a task token to be returned, and continues after a human response. Whether the implementation uses Step Functions or another durable workflow engine, the important properties are durable waiting, expiring tokens, authenticated callbacks, replay-safe transitions, and observable failure.

Prevent time-of-check/time-of-use failures

A correct proposal can become unsafe while it waits. Inventory changes, an account is suspended, the release artifact is replaced, or another operator has already applied the change. The executor must recheck conditions that can change.

For a deployment, bind approval to an immutable artifact digest, environment, configuration revision, and change ticket. For a payment, bind it to payee, amount, currency, source account, invoice identifier, and idempotency key. For deletion, bind it to the exact resource version and show dependent objects. If any bound field changes, fail closed and request approval again.

Keep the approval window short enough for the underlying assumptions. Expiration is a safety feature, not an inconvenience. A two-day-old approval for a production change usually says little about the current release candidate.

The reviewer experience determines control quality

“Approve / Deny” without context invites mistakes. The review surface should answer five questions at a glance:

  • What exact action will execute?
  • Which system, tenant, account, and environment will it affect?
  • What values or records will change?
  • Why was the action proposed, and which evidence supports it?
  • Can it be reversed, and what happens if it fails halfway?

Show a human-readable summary and the consequential structured fields. A production change screen should display the artifact digest and diff, not an editable model summary. A financial approval should make amount and payee visually prominent. A deletion review should list the resource and dependencies. Avoid presenting untrusted retrieved text as if it were a trusted system instruction.

The reviewer also needs meaningful choices. Deny should optionally capture a reason. “Edit” should generate a new proposal rather than mutating the approved object. “Escalate” should route the request to a different role. High-risk approvals may require step-up authentication or two different people, especially when the requester would otherwise approve their own proposal.

Approval policy belongs outside the prompt

A system prompt may tell the agent to ask before deleting records, but untrusted content can compete with that instruction and model behavior can vary. Enforce the rule where tools are dispatched. The model may propose customer.delete; the dispatcher marks that operation Tier 3; the executor accepts it only with a valid approval artifact issued for that proposal.

Useful deterministic policy inputs include:

  • authenticated requester and tenant;
  • tool and operation name;
  • target environment and resource sensitivity;
  • amount, count, geographic scope, and blast radius;
  • whether the operation is externally visible or destructive;
  • reversibility and rollback availability;
  • normal business hours or emergency mode;
  • the requester’s relationship to the required approver.

Use the model to help explain the proposal, not to waive the control. Teams defining these boundaries can connect approval design to broader cloud security and governance requirements such as identity, logging, separation of duties, and change management.

Audit the decision and the execution

An audit trail must reconstruct both what was authorized and what happened. Record the proposal identifier, tool and arguments, parameter digest, requester, reviewer, decision, timestamps, policy version, risk tier, reason, execution identity, target transaction identifier, outcome, and correlation ID. Record the model, prompt, and tool-schema versions needed to reproduce the route without logging sensitive prompt contents indiscriminately.

Protect audit records from the agent and the normal executor. Centralize them with retention and access appropriate to the action. Alert on unusual patterns: repeated denials, approval spikes, self-approval attempts, expired proposals being replayed, frequent policy overrides, or one reviewer approving at an implausible rate.

Amazon Bedrock agent tracing can expose orchestration and action information for diagnosis, but platform trace is only one part of the record. The target system’s own audit event and the application approval decision should share a correlation identifier.

Failure modes to test before launch

  • Rubber-stamp fatigue: too many low-risk prompts train reviewers to click without reading.
  • Approval after execution: the agent or tool has another path to the privileged API.
  • Mutable proposals: parameters can change after the decision.
  • Weak reviewer identity: an approval link functions as an unbounded bearer credential or leaks into logs.
  • Self-approval: the requester satisfies a control intended to create separation of duties.
  • Replay: an approved token performs the operation more than once.
  • Stale approval: execution ignores changed prerequisites or expired context.
  • Ambiguous outcome: a timeout causes an unsafe automatic retry.
  • Shadow tools: a broad “run command” or generic API tool bypasses the reviewed operation catalog.
  • Incomplete audit: the approval log and target transaction cannot be linked.

Exercise denial, cancellation, expiration, duplicate clicks, reviewer removal, policy changes, target outages, partial failures, and emergency shutdown. A kill switch should disable the privileged executor or individual operation without requiring a prompt update.

Measure whether human review is working

Track approval rate, denial reasons, time to decision, expiration rate, edits that create replacement proposals, execution failures after approval, duplicate-prevention events, and reviewer workload. Sample approved results for correctness. A 100 percent approval rate may mean the agent is excellent, but it may also mean the gate adds no scrutiny.

Use evidence to reduce unnecessary review. A low-impact action that passes policy for months may move to automated execution with notifications and sampling. An action with frequent edits or denials needs a better proposal, stronger validation, or a higher tier. The control should evolve without silently relaxing the business rule.

Recommended approach

Begin with a catalog of agent capabilities. Mark which are read-only, externally visible, financial, destructive, security-sensitive, or production-changing. Remove operations that cannot be bounded safely. Give the remaining operations typed schemas, service-side authorization, and risk tiers. Implement the approval workflow as durable application state, then test it without an AI model by submitting known proposals directly.

Only after that control path works should a model be allowed to propose actions. This ordering keeps the safety property independent of prompt quality. Bluegrass Cloud can help teams design and implement the surrounding AI solution, including governed tool interfaces, approval states, and operational ownership. To discuss a specific workflow, start with the action, affected systems, and decision owners.

Sources and further reading