AI Guides › Playbooks

The Judgement Ledger: Turning Your Expertise into a SKILL.md File

By Nigel Guy · 7 min read

Most people who try to "teach" an AI agent their job write a job description. "You are a senior test analyst with 12 years' experience. Be thorough." It feels like a handover. It isn't. The model already knows what a test analyst is; what it doesn't know is the three things you check before anything else, the question that always exposes the gap, and the mistake you made in 2019 and never made again. A title is theatre. The judgement is the bit that transfers.

The rule: capture decisions, not credentials. If a line in your skill file wouldn't change what the agent does next, delete it.

What a skill file actually is

A skill is a folder with a file called SKILL.md in it. The file opens with a short block of YAML frontmatter holding a name and a description, followed by plain Markdown instructions. Anthropic's Claude products use it, and the format is published as an open standard at agentskills.io; OpenAI's documentation says Codex and the ChatGPT desktop app read the same SKILL.md format too.

Three details decide whether yours gets used:

Detail What the docs say Why it matters to you
description Required, up to 1,024 characters. Should say what the skill does and when to use it, written in the third person It is the only part loaded at the start. If it's vague ("Helps with reviews"), the agent never picks the skill up
name Lowercase letters, numbers and hyphens, up to 64 characters; Anthropic also bars the words "anthropic" and "claude" A bad name fails validation or upload
Body length Keep SKILL.md under 500 lines; move long reference material into separate files one level down The whole body loads once the skill triggers, so padding costs context on every use

Where it lives depends on the tool. In Claude Code, a personal skill goes in ~/.claude/skills/<skill-name>/SKILL.md and a project skill in .claude/skills/<skill-name>/SKILL.md; you can also call it by typing /skill-name. In the Claude apps, you turn on Settings > Capabilities > Code execution and file creation, then upload a ZIP of the folder under Customize > Skills, using + > Create skill > Upload a skill. Anthropic's help centre lists skills on Free, Pro, Max, Team and Enterprise plans at time of writing; on Team and Enterprise an owner has to switch them on first.

The Judgement Ledger

Before you write the file, you need the raw material. The Judgement Ledger is a four-column table you fill in from real work, not memory of what you think you do.

Column What goes in it Test for keeping it
Trigger The situation you were in ("new requirements doc lands from a client") Would a stranger recognise this situation from the wording?
Move What you actually did first, second, third Is it an action, not an attitude? "Be careful" fails; "check every acceptance criterion has a measurable outcome" passes
Question The question you asked that caught something Has it caught a real problem at least once?
Scar A mistake you've seen or made, and what you do now instead Can you name the fix, not just the failure?

How to fill it:

  1. Pick one recurring task. One skill, one job. "Reviewing requirements" works; "everything I know about QA" doesn't.
  2. Work three real cases. Take three past examples, ideally one easy, one messy, one that went wrong. For each, write down every move and question in the order it happened.
  3. Strike anything the model already knows. Anthropic's own authoring guidance assumes the model is already capable and asks whether each line earns its token cost. Generic advice ("read carefully", "think about edge cases") goes.
  4. Keep what repeats. A move that appears in all three cases is a step. One that appears once is an edge case, which belongs in a short "watch for" list or nowhere.
  5. Write one sign-off example. A concrete piece of work you would approve, with a sentence on why. Examples carry standards better than adjectives.

Then turn the ledger into the file.

The skill file scaffold

Fill in the square brackets from your ledger. Leave out any section you have nothing real for.

---
name: [lowercase-hyphenated-name, e.g. reviewing-requirements]
description: [Third person. What it does + when to use it. e.g. "Reviews software requirements documents for testability, gaps and hidden risk. Use when the user shares a requirements doc, user story or acceptance criteria for review."]
---

# [Plain title of the task]

## Context
[One or two lines of domain context the model would not otherwise have, e.g. "Work is for regulated fintech clients; audit trail requirements are non-negotiable." No job titles or years of experience.]

## Steps — follow in order
1. [First move from your ledger, written as an action]
2. [Second move]
3. [Third move]

## Questions to raise every time
- [Question from your ledger that catches a common gap]
- [Question that exposes hidden risk]

## A result I would sign off
[One concrete example, plus one sentence on why it passes.]

## Known traps
- [Scar: the mistake] → [What to do instead]

## Output
[Exact format: e.g. "A table with columns Issue | Severity (High/Med/Low) | Suggested fix, then a three-line summary."]

## If information is missing
Do not assume. List what is missing and ask before reviewing.

## Before answering
Check that every step above was applied and every question was considered. Say which ones did not apply and why.

Fill in: the name, description and every bracket from your own ledger.

If you'd rather not start from a blank page, have a model interview you to fill the ledger. Paste this into any capable chat assistant:

You are an interviewer helping a practitioner capture their working judgement so an AI agent can reuse it.

Task: [THE_RECURRING_TASK, e.g. reviewing supplier contracts]
My background in one line: [YOUR_DOMAIN]
Three real cases I can draw on (anonymised): [CASE_1], [CASE_2], [CASE_3]

Process:
1. Ask me one question at a time, no more than 12 in total.
2. For each case, find out what I did first, what I checked next, which question uncovered a problem, and what I would do differently now.
3. Push back when I give an attitude ("I'm thorough") instead of an action; ask what I physically do.
4. Do not invent steps, examples or mistakes. If I haven't said it, it doesn't go in.

Output, once I say "done":
- A table with columns Trigger | Move | Question | Scar.
- Mark any row supported by only one case as "edge case".
- Then a draft SKILL.md using the sections: Context, Steps, Questions, Sign-off example, Known traps, Output, If information is missing.

Before you hand it over, check: is every row traceable to something I said? Remove anything that isn't.

Fill in: the task, your domain in a line, and three anonymised past cases.

Worked example (hypothetical)

Priya, a fictional compliance test analyst, keeps pasting the same review checklist into chat. She runs the interview prompt on three past requirements reviews. Her ledger shows that in all three she first checked whether each acceptance criterion could actually fail a test, then asked "who signs this off, and what evidence will they want?" Her scar: she once approved a story where "the system logs the change" didn't say which fields were logged, and the audit failed.

She cuts the career paragraph from her first draft. The final SKILL.md is about 40 lines: three steps, two questions, one sign-off example, one known trap ("'logs the change' → name the fields, retention period and who can read the log"), and a table output format. She tests it on a fourth requirements doc it hasn't seen, notices it skipped the sign-off question, and moves that question into Step 2 so it can't be missed.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo