AI Guides › Playbooks

The Context, Role and Check Card for Working With AI Agents

By Nigel Guy · 8 min read

Most people hand an agent a one-line request, wait, and then skim the result for anything that looks wrong. That feels efficient because the agent writes confidently and fast. It fails because the agent has none of the background you carry in your head, no defined job, and nothing it can test its own work against, so you become the only quality control and you are checking by eye.

The rule: write the context down, give the agent a role with edges, and never accept output that has not passed a check someone other than the author can see.

One caveat first. Anthropic has not, as far as we could verify, published a single document titled as a guide to running human-agent teams. What it has published is spread across its Claude Code documentation and engineering blog, and the three rules below are our reading of those sources. Where we cite a practice, the source is listed at the end.

The mechanism: the Context, Role and Check Card

It is one page you fill in before delegating anything that will take an agent more than a few minutes. Three boxes, in this order.

Box Question it answers Where it lives
Context What does the agent need to know that it cannot guess? A file the agent reads at the start
Role What is the job, and where does it stop? The brief, or a saved agent definition
Check What result proves the work is done, and who looks at the evidence? A test, a script, a second reviewer, and you

Box 1 — Context: why does it matter what is written down?

An agent does not remember your last session and cannot read your mind. Anthropic's Claude Code best-practices page says its advice rests on one constraint: the context window fills up fast, and performance degrades as it fills. So context has two jobs, getting the right facts in and keeping the wrong ones out.

What to write down, in a persistent file such as CLAUDE.md (a file Claude Code reads at the start of every conversation):

What to leave out: anything the agent can work out by reading the material itself, and anything that changes weekly. The same page offers a useful test for each line: would removing it cause a mistake? If not, cut it. A bloated file means the rules that matter get lost, which is the opposite of what you wanted.

Long-running work needs the same habit across sessions. Anthropic's engineering post on long-running agents describes a progress file the agent keeps updated, a structured list of required features with a pass or fail marker, and git history the next session reads before it starts. The principle is plain: the next session starts blank, so the state has to be on disk.

Box 2 — Role: what is the difference between a role and a task list?

A task list says what to do. A role says what the job is, what good looks like, what is out of bounds, and what to hand back. The difference shows up when something unexpected happens: a task list has no answer, a role does.

Anthropic's write-up of its multi-agent research system names four things a delegated brief needs: an objective, an output format, guidance on tools and sources, and clear task boundaries. It reports that without detailed task descriptions, agents duplicate work, leave gaps, or fail to find what they need. It also found agents struggle to judge how much effort a task deserves, so scaling rules (a simple lookup gets a little effort, a broad comparison gets more) were written into the prompts.

In Claude Code you can save a role as a subagent: a Markdown file in .claude/agents/ (this project) or ~/.claude/agents/ (all your projects), with name and description in the frontmatter, and optionally tools and model. The tools field restricts what the role may touch. A reviewer role that only has Read, Grep, Glob cannot quietly edit the thing it is reviewing. Per the current docs, a subagent starts in a fresh context and does not see your conversation, so the brief you give it is all it has. Check the sub-agents page for current field names, as these change.

Here is a brief that fills the Role box. Replace the bracketed items.

You are a [ROLE, e.g. careful research assistant] working for [WHO YOU WORK FOR].

Goal: [ONE-SENTENCE OUTCOME]. A good result means [WHAT "DONE" LOOKS LIKE].

Context you must use: [FILES, LINKS OR NOTES]. Decisions already made, do not reopen: [LIST].

Scope: you may [ALLOWED ACTIONS]. You must not [OUT-OF-BOUNDS ACTIONS]. If a task needs something outside this scope, stop and ask.

Steps:
1. Restate the goal and list any inputs you are missing. Ask me for them rather than guessing.
2. Do the work in the order that [SENSIBLE ORDER].
3. Check your own work against [CHECK, e.g. these test cases or this checklist].

Output format: [FORMAT, e.g. a table with columns X, Y, Z], followed by a short list of anything you were unsure about and any assumption you made.

Before you answer, confirm that every claim you made is traceable to a source I gave you or that you cite, and say plainly where it is not.

Fill in: the role, scope and format. Leave step 1 in; it is what stops the agent guessing.

Box 3 — Check: how do you stop trusting output blindly?

The best-practices page is blunt on this. Claude stops when the work looks done, and without a check it can run, "looks done" is the only signal, so you become the verification loop. Its listed failure pattern, the trust-then-verify gap, ends with: if you cannot verify it, do not ship it.

Build the check in layers, cheapest first:

  1. A pass or fail the agent can run itself. A test suite, a build exit code, a script that compares output to a known example. Put the criteria in the brief.
  2. Evidence, not assertions. Ask for the command it ran and what came back, or a screenshot, rather than "all done".
  3. A reviewer that did not write it. The docs recommend a fresh subagent or session reviewing the result, because it sees the result and your criteria, not the reasoning that produced it. Tell it to report only gaps affecting correctness or stated requirements; a reviewer asked to find problems will usually find some, and chasing all of them causes over-engineering.
  4. A human on the edge cases. Anthropic's research-system post says automated evaluation missed problems that people testing the agent caught, including a bias towards poorly sourced pages. Your own sampling is not optional.

For reviews, this prompt does the job:

You are an independent reviewer. You did not produce this work and you have not seen the reasoning behind it.

Review [WORK TO CHECK] against [REQUIREMENTS OR PLAN].

Report only gaps that affect correctness or a stated requirement. Ignore style preferences. For each gap give: what is wrong, where, the evidence, and a suggested fix. If you find no gaps, say so and list what you checked.

If you lack something needed to judge, ask me; do not assume.
Before answering, remove any finding you cannot support with evidence.

A worked example (hypothetical)

A small bakery owner, Priya, wants an agent to reconcile a month of supplier invoices against her spreadsheet.

The first run flags the disputed March fee again. That tells her the context note was ambiguous, so she fixes the note, not the agent.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo