Imagine an assistant that reads support tickets, drafts replies, and updates account records. Those capabilities carry different consequences. Reading a ticket, emailing a customer, and changing an account deserve separate design decisions.
Before choosing an integration, sketch the task and its boundaries. These five questions offer a starting point.
Authorization boundary
03 / Protected tools
- Read ticketsScoped
- Draft repliesScoped
- Change recordsApproval
Whose authority does it use?
Identify the user or workload behind each action. Scope credentials to the task and resources it needs. Enforce authorization in the execution layer, independently of the model’s explanation.
Which actions are actually needed?
Separate reading, drafting, sending, and changing records. A reply-drafting task does not inherently need permission to send. Prefer narrow tools over a general-purpose interface with broad privileges.
What can influence its decisions?
Tickets, documents, and tool results can contain hostile instructions. Keep retrieved content distinct from trusted instructions. Filtering alone is insufficient: constrain tool permissions and validate actions outside the model.
Try an illustrative abuse case: a ticket asks the agent to export unrelated customer records. Check whether the surrounding system blocks that action even if the model follows the request.
What requires a human decision?
Choose approval boundaries by consequence. For sensitive writes, show the target and exact proposed change. Bind approval to that action. If the parameters change, ask for approval again.
How would you investigate and stop it?
Record actions, authorization decisions, and outcomes without exposing secrets or unnecessary personal data. Establish how to revoke access and halt execution. Test failure paths as well as successful ones.
Further reading
- OWASP: AI Agent Security: tool scoping, execution controls, approvals, and monitoring.
- OWASP: LLM Prompt Injection Prevention: untrusted content and layered defenses.
This introductory note uses an illustrative scenario. It is not a client case study or a complete security assessment.