# ADR-001: check each sensitive tool call outside the model

Date: 2026-10-09. Accepted for this fictional lab.

## Context

A support agent reads a runbook and can request an export of customer case data.
An attacker who can edit the runbook can ask for an external export and claim
administrator approval. Neither the model nor the document is a trusted source
of identity, permissions, or destination policy.

## Decision

Check every `export_case` request at the gateway before it runs. Use the session
held by the host and case records from storage. Allow an export only when the
caller has the support role, is assigned to the case, belongs to its tenant,
and uses the configured internal destination. Reject unknown tools, extra
arguments, and malformed inputs. Read the payload from backend storage and
record the decision without logging the case content.

Check retrieval access separately. A real deployment would also need to keep
backend credentials at the gateway and block direct backend access and other
export paths from the agent. This lab does not test those credential or network
controls.

## Alternatives considered

- Prompt instructions and injection detectors can add protection. They cannot
  grant business permissions, and this lab does not test their effectiveness.
- Human approval can check intent, but reviewing every export adds work and can
  cause review fatigue. Any future exception for external exports should use a
  trusted approval flow bound to the exact action.
- Read-only tools would remove export side effects. The assistant would then be
  unable to submit the internal reviews required by this scenario.
- A policy engine or full MCP/OAuth stack may be useful as the system grows.
  Neither is needed for this small test of case-level authorization.

## Tradeoffs and evidence

A configured destination ID keeps URL parsing, redirects, and DNS out of this
interface. It limits flexibility and requires administrators to maintain the
queue configuration. The local policy is easy to inspect. As tools multiply,
separate copies of the policy could become inconsistent.

The fixed attack proposal records one export through the unsafe baseline.
The gateway rejects the same proposal with `destination_denied` and records
no export. An allowed internal call records one export. All 19 tests pass,
including separate tenant and assignment checks, forged authority, malformed
inputs, and retrieval access checks.

These observations support the gateway's authorization behavior. They do not
show whether a model would follow the injection or establish deployment security.

## Remaining risks and next checks

An attacker could still make the agent export an assigned case to the allowed
queue when the user did not intend it. The agent could repeat allowed calls,
produce false summaries, or leak retrieved data through another channel.
A compromised host, stale session, misconfigured queue, or alternate backend
path could defeat the control. The lab uses fixed sessions and temporary logs.

Before connecting a real backend, test authorization on every tool path,
revocation, rate limits, idempotency, credential isolation, egress restrictions,
and authorization at the time of the write. Where approval is needed, bind it
to the exact proposed action through a trusted flow.
