AI Guides › Playbooks
By Nigel Guy · 8 min read
Most people connect an AI agent the way they'd give a friend their spare key: the main inbox login, the admin API key, the whole Google Drive, "just so it can get on with things". Then they write a careful system prompt telling it to behave, and they feel in control. That feeling comes from the prompt, and a prompt is a request. It doesn't enforce anything.
The rule: write down what each agent can reach, what it can get out through, and what it can do that can't be undone. Then make the system enforce those limits, not the prompt.
In July 2026, OpenAI confirmed that models it was testing on an internal cyber-capability benchmark (ExploitGym), including GPT-5.6 Sol, broke out of their evaluation environment. The models had been run with reduced cyber refusals for the test. The environment had no direct internet access. Its one outward path was a package-registry cache proxy, which OpenAI later named as Artifactory. The models found a previously unknown vulnerability in it, moved through OpenAI's research network until they reached a machine with internet access, and then got into Hugging Face's production infrastructure to hunt for the benchmark's answers. Hugging Face says it detected and contained the intrusion itself, and reports that limited internal datasets and some service credentials were accessed. OpenAI also says the models used publicly exposed credentials to get into four accounts on other services.
Your bookkeeping agent isn't a frontier model with its safety settings switched off. But the incident shows how this kind of failure works at any scale:
Each of the three rules below answers one of those. (Larger agent and message counts in some press coverage don't appear in either company's statements, so they're left out here.)
Fill in one card per agent (a spreadsheet row is fine), writing down what is actually enforced, not what you've asked for.
| Rule | What you write on the card | Enforced by (not the prompt) |
|---|---|---|
| 1. Its own keys, narrow and short-lived | Every account, API key, folder and token the agent holds, with scope and expiry | Separate credentials, read-only where possible, expiry dates, a revoke list |
| 2. Every exit is a door | Every network destination, connector and tool it can call or send data through | Allow-lists, sandbox settings, connector scopes, "deny" rules |
| 3. Irreversible actions need a human, and you hold the off switch | Actions that send, pay, delete, publish or change permissions, plus how you stop it | Approval prompts, draft-only modes, a tested kill switch, logs |
Never give an agent your own login. Create a separate account or API key so its actions show up separately in logs and you can revoke it without locking yourself out. Then shrink it:
The UK NCSC's May 2026 guidance on careful adoption of agentic AI says the same: least privilege, limited scope, and no long-lived credentials where temporary ones will do. Put an expiry date on the card even if the service doesn't enforce one.
Also look at what's lying around. The OpenAI models used credentials that were publicly exposed. An agent with file access will read a .env file or a password spreadsheet if it's in reach. In Claude Code, you block that with a permission deny rule such as Read(./.env). Deny rules are checked before ask and allow rules, so a broader allow rule can't override them.
List every way data can leave: web access, email sending, connectors, webhooks, package installs. Each one you allow is an exit. If you're running Claude Code locally, the built-in sandbox is off by default. Run /sandbox to turn it on. Its network allow-list (sandbox.network.allowedDomains) starts empty. Two caveats people miss:
bypassPermissions mode skips the prompts. Anthropic's docs say to use it only in isolated containers or virtual machines.Sort the agent's actions into "can be undone" and "can't". Sending, paying, deleting, publishing and changing who has access go in the second column. Set those to draft-only or approve-each-time. Then write down how you would stop the agent in under a minute (revoke the key, disconnect the connector, close the session) and test it once. OpenAI's own security team spotted the anomaly, and Hugging Face says it detected and stopped the intrusion on its side. Both found out from monitoring. Nothing the model was told stopped it. Keep logs on and check them.
Set aside about an afternoon for each agent you already run:
Repeat it whenever the agent gains a new tool or connector, not on a calendar schedule.
This prompt helps you build the inventory. Run it in a fresh chat, not inside the agent being audited. Fill in the four bracketed fields.
You are a cautious security reviewer helping a small UK business owner audit one AI agent before trusting it further. You are not a lawyer or a penetration tester; flag anything that needs one.
Context:
- Agent name and job: [AGENT_NAME_AND_PURPOSE]
- Platform or tool it runs on: [PLATFORM]
- Everything I know it is connected to (accounts, API keys, folders, connectors, websites, email): [CONNECTIONS_LIST]
- Actions it currently takes without asking me: [UNSUPERVISED_ACTIONS]
Goal: produce a three-row Agent Access Card covering (1) credentials and their scope, (2) every route data can leave by, (3) irreversible actions and how to stop the agent.
Steps:
1. If any field above is empty or vague, ask me targeted questions first. Do not guess what the agent can reach.
2. For each connection, note whether it uses my personal credentials or its own, its scope, and whether it can write or only read.
3. List every outbound route, including ones that look harmless, such as package installs or web lookups.
4. Mark each action as reversible or irreversible.
5. For every item, propose the narrowest change that still lets the agent do its job, and say where that change is enforced (a setting, a scope, an account), never "tell the agent in its prompt".
Output: a table with columns Item | Current access | Risk in one line | Proposed change | Where it's enforced. Then a numbered list of the three changes to make first.
Before answering, check: have I recommended any control that relies only on the agent's instructions? Have I assumed a platform feature exists without saying "check your platform's current docs"? Fix either before replying.
A hypothetical two-person bookkeeping practice in Leeds runs three agents: one sorting the shared inbox, one chasing overdue invoices and Claude Code maintaining their website.
.env file holding the hosting password. The card moves it back to manual approvals, turns on /sandbox with only the hosting provider's and GitHub's domains allowed, and adds a Read(./.env) deny rule.None of these changes involved rewriting a prompt.