AI Guides › Playbooks
By Nigel Guy · 7 min read
Most people who try to "teach" an AI agent their job write a job description. "You are a senior test analyst with 12 years' experience. Be thorough." It feels like a handover. It isn't. The model already knows what a test analyst is; what it doesn't know is the three things you check before anything else, the question that always exposes the gap, and the mistake you made in 2019 and never made again. A title is theatre. The judgement is the bit that transfers.
The rule: capture decisions, not credentials. If a line in your skill file wouldn't change what the agent does next, delete it.
A skill is a folder with a file called SKILL.md in it. The file opens with a short block of YAML frontmatter holding a name and a description, followed by plain Markdown instructions. Anthropic's Claude products use it, and the format is published as an open standard at agentskills.io; OpenAI's documentation says Codex and the ChatGPT desktop app read the same SKILL.md format too.
Three details decide whether yours gets used:
| Detail | What the docs say | Why it matters to you |
|---|---|---|
description |
Required, up to 1,024 characters. Should say what the skill does and when to use it, written in the third person | It is the only part loaded at the start. If it's vague ("Helps with reviews"), the agent never picks the skill up |
name |
Lowercase letters, numbers and hyphens, up to 64 characters; Anthropic also bars the words "anthropic" and "claude" | A bad name fails validation or upload |
| Body length | Keep SKILL.md under 500 lines; move long reference material into separate files one level down |
The whole body loads once the skill triggers, so padding costs context on every use |
Where it lives depends on the tool. In Claude Code, a personal skill goes in ~/.claude/skills/<skill-name>/SKILL.md and a project skill in .claude/skills/<skill-name>/SKILL.md; you can also call it by typing /skill-name. In the Claude apps, you turn on Settings > Capabilities > Code execution and file creation, then upload a ZIP of the folder under Customize > Skills, using + > Create skill > Upload a skill. Anthropic's help centre lists skills on Free, Pro, Max, Team and Enterprise plans at time of writing; on Team and Enterprise an owner has to switch them on first.
Before you write the file, you need the raw material. The Judgement Ledger is a four-column table you fill in from real work, not memory of what you think you do.
| Column | What goes in it | Test for keeping it |
|---|---|---|
| Trigger | The situation you were in ("new requirements doc lands from a client") | Would a stranger recognise this situation from the wording? |
| Move | What you actually did first, second, third | Is it an action, not an attitude? "Be careful" fails; "check every acceptance criterion has a measurable outcome" passes |
| Question | The question you asked that caught something | Has it caught a real problem at least once? |
| Scar | A mistake you've seen or made, and what you do now instead | Can you name the fix, not just the failure? |
How to fill it:
Then turn the ledger into the file.
Fill in the square brackets from your ledger. Leave out any section you have nothing real for.
---
name: [lowercase-hyphenated-name, e.g. reviewing-requirements]
description: [Third person. What it does + when to use it. e.g. "Reviews software requirements documents for testability, gaps and hidden risk. Use when the user shares a requirements doc, user story or acceptance criteria for review."]
---
# [Plain title of the task]
## Context
[One or two lines of domain context the model would not otherwise have, e.g. "Work is for regulated fintech clients; audit trail requirements are non-negotiable." No job titles or years of experience.]
## Steps — follow in order
1. [First move from your ledger, written as an action]
2. [Second move]
3. [Third move]
## Questions to raise every time
- [Question from your ledger that catches a common gap]
- [Question that exposes hidden risk]
## A result I would sign off
[One concrete example, plus one sentence on why it passes.]
## Known traps
- [Scar: the mistake] → [What to do instead]
## Output
[Exact format: e.g. "A table with columns Issue | Severity (High/Med/Low) | Suggested fix, then a three-line summary."]
## If information is missing
Do not assume. List what is missing and ask before reviewing.
## Before answering
Check that every step above was applied and every question was considered. Say which ones did not apply and why.
Fill in: the name, description and every bracket from your own ledger.
If you'd rather not start from a blank page, have a model interview you to fill the ledger. Paste this into any capable chat assistant:
You are an interviewer helping a practitioner capture their working judgement so an AI agent can reuse it.
Task: [THE_RECURRING_TASK, e.g. reviewing supplier contracts]
My background in one line: [YOUR_DOMAIN]
Three real cases I can draw on (anonymised): [CASE_1], [CASE_2], [CASE_3]
Process:
1. Ask me one question at a time, no more than 12 in total.
2. For each case, find out what I did first, what I checked next, which question uncovered a problem, and what I would do differently now.
3. Push back when I give an attitude ("I'm thorough") instead of an action; ask what I physically do.
4. Do not invent steps, examples or mistakes. If I haven't said it, it doesn't go in.
Output, once I say "done":
- A table with columns Trigger | Move | Question | Scar.
- Mark any row supported by only one case as "edge case".
- Then a draft SKILL.md using the sections: Context, Steps, Questions, Sign-off example, Known traps, Output, If information is missing.
Before you hand it over, check: is every row traceable to something I said? Remove anything that isn't.
Fill in: the task, your domain in a line, and three anonymised past cases.
Priya, a fictional compliance test analyst, keeps pasting the same review checklist into chat. She runs the interview prompt on three past requirements reviews. Her ledger shows that in all three she first checked whether each acceptance criterion could actually fail a test, then asked "who signs this off, and what evidence will they want?" Her scar: she once approved a story where "the system logs the change" didn't say which fields were logged, and the audit failed.
She cuts the career paragraph from her first draft. The final SKILL.md is about 40 lines: three steps, two questions, one sign-off example, one known trap ("'logs the change' → name the fields, retention period and who can read the log"), and a table output format. She tests it on a fourth requirements doc it hasn't seen, notices it skipped the sign-off question, and moves that question into Step 2 so it can't be missed.