AI Guides › Playbooks
By Nigel Guy · 7 min read
Most people connect an agent the way they accept cookies: click through, grant everything, and assume the model will use the access sensibly. Then they write "never delete anything" in the instructions and call that a safety measure. It feels fine because the agent behaves well nine hundred times in a row. The thousandth time it misreads a file, or obeys text hidden in a web page, and the only thing between it and your inbox, repo or accounts is a sentence it was free to ignore.
The rule: decide what an agent can touch before it starts, enforce that outside the model, give each job its own keys, and plan for the day it goes wrong.
Before you connect an agent to anything real, fill in one card per job. It has four lines, and each line maps to one of the three rules plus the plan for failure.
| Line | Question | Example answer |
|---|---|---|
| 1. Reach | What can it read, change, send or spend? | Read one folder; draft emails, never send |
| 2. Enforcement | Which non-model control holds that line? | Permission rule, sandbox, read-only token |
| 3. Keys | Which credential is used, and is it used for anything else? | A dedicated token for this job only |
| 4. Undo | If it goes wrong, how do you stop it and recover? | Revoke the token; restore from backup |
If you cannot fill in a line, the agent is not ready to be connected. That is the whole test.
An instruction in a prompt is a request. A permission rule, a sandbox or a token scope is a wall. The OWASP guidance on "excessive agency" names three root causes: too much functionality, too many permissions and too much autonomy. Its recommended fixes are the structural ones: limit the tools an agent has, avoid open-ended tools such as raw shell commands or arbitrary URL fetching, apply least privilege on the downstream systems, and require human approval for high-impact actions. Note that none of that is "write a better prompt".
Take Claude Code as a concrete example, since its documentation is explicit about this:
Bash, removes the tool from Claude's context entirely. A scoped rule such as Bash(rm *) leaves the tool available and blocks matching calls.Read deny rule for the path, for example Read(./.env) or Read(./secrets/**).PreToolUse hook runs before a tool call and can block it; the documentation's own example exits with code 2 to deny a command containing rm -rf. The point of a hook is that it is deterministic: it runs whether or not the model "remembers" the rule./sandbox or by setting sandbox.enabled to true in a settings file. It restricts which files and network domains shell commands can reach, and it is enforced by the operating system.Read the catches too, because they are where beginners get hurt. The documentation says a Bash deny rule does not match the same program run by path or inside sh -c, and that Read and Edit deny rules do not stop a Python or Node script that opens files itself. The sandbox covers shell commands only: file tools, MCP servers and hooks run outside it. So a deny rule alone is not a wall; pair it with the sandbox when the restriction must hold. Other agent products have their own equivalents, so look for the same four things: permission rules, an approval step, an isolated environment and scoped tokens. This area changes often, so check the current page for your tool rather than trusting a remembered menu.
The shortcut is to paste your own admin credentials into the agent because they already work. Do not. Whatever the agent holds, anyone who can steer it (through a poisoned web page, a hostile email or a compromised tool) can use.
For each job:
A separate key also gives you a clean log. When something changes in the account, you can tell which job did it.
The Undo line of the card is not pessimism; it is what makes the other three affordable. Before the first run:
Say you run a small online shop and want an agent to tidy supplier invoices in a shared folder and draft chasing emails.
| Line | Card entry |
|---|---|
| Reach | Read one invoices folder; write a summary file in a separate output folder; draft emails only |
| Enforcement | Run in a sandbox or container with only that folder mounted; no send permission on the mailbox connection; deny rule on the finance exports folder |
| Keys | A new read-only token for the folder service; a mailbox connection limited to creating drafts; neither used elsewhere |
| Undo | Folder backed up nightly; revoke both tokens from the provider's settings; drafts reviewed by you before sending |
Notice what is absent: bank access, the shop admin login, and any "send" permission. If a later job needs them, it gets its own card and its own keys.
A prompt can help you fill in the card honestly. Paste in your planned setup and let it argue with you:
You are a cautious security reviewer helping a non-specialist. I plan to connect an AI agent to the following: [TOOLS_AND_ACCOUNTS]. The job is: [JOB_DESCRIPTION]. The access I intend to grant is: [PLANNED_PERMISSIONS].
Goal: fill in a four-line Agent Access Card (Reach, Enforcement, Keys, Undo) and flag anything that is broader than the job needs.
Steps:
1. List everything the agent could read, change, send or spend with this setup.
2. Mark each item as needed or not needed for the job.
3. For each not-needed item, suggest a narrower alternative.
4. For each remaining risk, say whether it is controlled by a rule outside the model (permission, sandbox, token scope) or only by instruction, and treat instruction-only as uncontrolled.
5. Describe how I would stop the agent and recover.
Output: the card as a table, then a short list of changes to make before the first run.
Constraints: if you lack information about the tool or platform, ask me instead of guessing. Do not claim a setting exists unless you are sure; tell me to check the official documentation for it.
Before answering, check that every line of the card has a concrete, non-model control or is clearly marked as a gap.
Fill in the three bracketed items. Treat the answer as a draft to check against the vendor's current documentation.