AI Guides › Playbooks
By Nigel Guy · 7 min read
Most people meet "agents" by pasting a bigger and bigger request into one chat and hoping. It feels like progress because the answer gets longer, but one context window is now doing research, drafting, checking and formatting at once, and nobody is responsible for any of it. An agent team fixes that by giving each piece of the job an owner, an input, an output and a stopping point.
The rule: never create an agent without writing down what it receives, what it hands back, what it may touch, and when it stops.
In July 2026 Meta Superintelligence Labs released Muse Spark 1.1. Meta describes it as a multimodal reasoning model built for agentic tasks, trained to orchestrate multi-agent systems: as the main agent it gathers context, makes a plan and delegates to parallel subagents; as a subagent it sticks to its job and knows when to escalate back. Meta also says it can manage a context window of 1 million tokens, and it launched a public preview of the Meta Model API alongside it. It is available in "Thinking" mode in the Meta AI app and on meta.ai.
Those are Meta's own claims, taken from Meta's own announcement. I could not verify independent test results, and I could not confirm API pricing, so this guide gives none. Check Meta's developer pages before you plan a budget.
The durable part is the pattern, and it works with any capable model. Claude Code, for instance, has two built-in versions of it:
| Chatbot | Subagents | Agent team | |
|---|---|---|---|
| Who does the work | One model, one thread | Helpers that each get their own context and report back to the caller | Separate sessions that message each other and share a task list |
| Good for | Questions, drafts | Focused jobs where only the result matters | Work where the workers need to challenge each other |
| Cost | Lowest | Lower | Higher: each teammate is its own instance |
| Status in Claude Code | Default | Available | Experimental, off by default |
In Claude Code, agent teams are switched on by setting the environment variable CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS to 1 in your settings. Anthropic's documentation notes that teams use significantly more tokens than a single session and recommends starting with three to five teammates. It also lists known limits, among them that resuming a session does not restore in-process teammates. Check the current page, because experimental features move.
A Role Sheet is one table, written before any agent runs. One row per agent.
| Field | What you write | Why it matters |
|---|---|---|
| Role | One verb phrase: "Collect last week's figures" | Vague roles overlap |
| Input | Exactly what it receives, and from where | Agents don't inherit your chat history |
| Output | The format and length it hands back | The next agent needs something predictable |
| Allowed to touch | Files, tools, accounts: read-only unless stated | Limits the damage a bad run can do |
| Stop condition | Done when X, or escalate when Y | Stops loops and silent guessing |
| Checker | Who reviews it: another agent or you | Someone must be able to say no |
Then build in four steps.
Imagine a small online shop owner who wants a Monday brief. These figures and tools are invented for illustration.
| Role | Input | Output | Allowed to touch | Stop condition |
|---|---|---|---|---|
| Collector | Last week's sales export, support inbox summary, ad spend sheet | Three tidy tables, nothing interpreted | Read-only | Missing file: say so, do not estimate |
| Analyst | The three tables | Five bullet findings, each tied to a table row | None | Fewer than five real findings: report fewer |
| Drafter | Findings, last week's brief for tone | One-page brief, 300 words maximum | Draft file only | Brief written or a question asked |
| Checker | Tables, findings, brief | Pass or list of mismatches | None | Any number not traceable to a table |
The owner reads the checker's verdict first, then the brief. For the first month, the owner also spot-checks two numbers by hand.
In Claude Code, each row can be a subagent: a markdown file in .claude/agents/ (project) or ~/.claude/agents/ (all projects) with a name, a description, an optional tools allowlist such as Read, Grep, Glob, an optional model, and a system prompt. That is the Role Sheet turned into a file.
Lead prompt. Fill in the job, the roles and your limits.
You are the lead on a small agent team. Your job is to plan, delegate and merge, not to do every task yourself.
The job, in one checkable sentence: [JOB_SENTENCE]
What a good result looks like: [FORMAT_AND_LENGTH]
Roles available: [ROLE_LIST_WITH_INPUT_OUTPUT_AND_ALLOWED_TOOLS]
Hard limits: [WHAT_NOBODY_MAY_DO, E.G. SEND EMAILS OR EDIT LIVE FILES]
Steps:
1. If any role has no clear input or output, ask me before starting. Do not guess.
2. Give each helper only the context it needs, in writing, with its stop condition.
3. Collect the outputs and merge them. Where two outputs disagree, say so rather than choosing quietly.
4. Pass the merged result to the checker role before showing me.
Before answering, check: does every claim trace to a helper output, and did anyone step outside their limits? Report that check at the end in two lines.
Role prompt, one per agent.
You are the [ROLE_NAME] on a team producing [JOB_SENTENCE].
You will receive: [INPUT_DESCRIPTION].
Hand back: [OUTPUT_FORMAT_AND_LENGTH]. Nothing else.
You may use only: [ALLOWED_TOOLS_AND_FILES]. Treat everything else as off limits.
Stop when: [STOP_CONDITION]. If an input is missing or unclear, say exactly what is missing and stop; do not fill gaps with guesses.
Before you reply, check that every figure or claim in your output comes from your input, and mark anything you could not source as UNSOURCED.
You can name the owner of any sentence in the output. A missing input produces a stated gap, not a plausible invention. The checker has rejected something at least once in the first fortnight; a checker that never objects is decoration. And the team beats you doing the job alone on time or on quality, not merely on volume.