AI Guides › Workbench
By Nigel Guy · 6 min read
You ask Claude whether your plan is good, and it finds reasons to say yes. It feels like feedback and works like a mirror. One voice, one framing and one polite tone give you no way to see what a sceptic, a finance person or a customer would have said.
The rule: never take one answer to a decision. Make the model argue the question from five fixed positions, make them disagree before anyone summarises, and have a chair that must name what stays unresolved.
Andrej Karpathy published a repository called llm-council. Per its README, it sends your question to several different models through OpenRouter, shows their first answers, has each model rank the others' anonymised answers, then has a designated "Chairman" model write the final response. The default members at time of writing are GPT-5.1, Gemini 3.0 Pro, Claude Sonnet 4.5 and Grok 4, with Gemini as chair. Karpathy describes it as a weekend hack, provided as is, with no support promised.
The kit below is different: it imitates the shape (separate opinions, ranking, chair) inside one Claude conversation. I could not verify any official Karpathy version that runs "five advisers inside one Claude", so treat this as an adaptation, not his tool. One model playing five roles is cheaper and quicker than a real multi-model council, and also weaker, because every adviser shares the same training and the same blind spots.
| Option | What it does | Cost at time of writing | Best for | Catch |
|---|---|---|---|---|
| One-chat council (prompt A) | Five advisers and a chair in one reply | Included in your Claude plan, including Free | Fast checks on a decision | Same model behind every voice, so agreement is partly staged |
| Separate-chats council (prompt B) | Each adviser in a fresh chat, then a chair chat | Uses more of your plan's usage allowance | Decisions where independence matters | Slower, and you paste text between chats |
| Karpathy's repo | Real multiple models, anonymous ranking | Free code, but you pay OpenRouter for credits and need to run a Python backend and React frontend | Technical readers who want genuinely different models | Unsupported by its author; costs vary by model |
Claude's consumer plans at time of writing: Free at £0, Pro at $20 a month on monthly billing (about £15 at time of writing, check the £ price at checkout), Max from $100 a month (about £75 at time of writing, check at checkout). Plans, prices and limits change, so check claude.com/pricing.
Fixed seats stop the council drifting into five polite versions of the same view. These are mine, not Karpathy's. Swap any seat that does not fit your decision.
| Seat | Job |
|---|---|
| The Sceptic | Finds the likeliest way this fails |
| The Operator | Asks what has to happen on Monday, and who does it |
| The Numbers Person | Tests the money, time and evidence |
| The Customer | Speaks for the person on the receiving end |
| The Outsider | Questions the framing, and suggests a different option |
The chair is not a sixth opinion. Its job is to compare the five and decide what the reader should do next.
Fill in the decision, your context and your constraints. Paste real numbers; do not let it invent them.
You are running a decision council for me. You will play five advisers
in turn, then a neutral chair.
MY DECISION: [DECISION_AS_A_QUESTION]
CONTEXT: [BACKGROUND_FACTS_AND_NUMBERS_YOU_KNOW]
OPTIONS I AM CHOOSING BETWEEN: [OPTION_1], [OPTION_2], [DOING_NOTHING]
WHAT A GOOD OUTCOME LOOKS LIKE: [SUCCESS_MEASURE_AND_DEADLINE]
CONSTRAINTS: [BUDGET_TIME_PEOPLE_RULES]
Before you start, check the inputs. If anything above is missing or too
vague to judge, ask me up to five short questions and wait. Do not guess
facts about my situation.
Step 1. Write each adviser's view separately, in this order, without
referring to any other adviser's view:
- The Sceptic: the most likely way this fails.
- The Operator: what must happen in the first two weeks, and by whom.
- The Numbers Person: what the money, time and evidence say. Label every
figure as "from me" or "your assumption". Invent no statistics.
- The Customer: how the person affected would react, in plain words.
- The Outsider: whether I am asking the right question, and one option
I have not listed.
Each adviser gets 80 to 120 words, one clear recommendation, and one
thing that would change their mind.
Step 2. Rank the five views from most to least useful for my decision.
Give one line of reasoning each. Do not rank by who sounds most confident.
Step 3. As the chair, write:
- Where the advisers agree.
- Where they genuinely disagree, and what evidence would settle it.
- Your recommendation, with the main risk.
- The cheapest test I can run in seven days.
- What you could not judge from the information I gave you.
Self-check before you answer: did every adviser disagree with at least
one other on something real? Did any figure appear that I did not give
you? If either answer is wrong, fix it first.
Independence is the point of Karpathy's design, where models answer before seeing each other. To get closer, open five fresh chats, put one seat in each, then bring the answers to a sixth.
For each of the five chats, paste this with the seat filled in:
You are [SEAT_NAME] on a decision council: [SEAT_JOB]. Judge this
decision from that position only.
DECISION: [DECISION_AS_A_QUESTION]
CONTEXT: [BACKGROUND_FACTS_AND_NUMBERS_YOU_KNOW]
OPTIONS: [OPTION_1], [OPTION_2], [DOING_NOTHING]
If key facts are missing, ask me before answering. Give a recommendation
in 150 words or fewer, the single biggest risk from your seat, and what
evidence would change your view. Do not invent figures.
Then, in a sixth chat, strip the seat names off the five answers, label them Answer 1 to 5, and paste them under this:
You are the chair of a decision council. Below are five anonymous
answers to the same question. Decision: [DECISION_AS_A_QUESTION].
[PASTE_ANSWER_1_TO_5_HERE]
First, rank the answers by how well they use evidence, with one line of
reasoning each. Then state the points of agreement, the real
disagreements, your recommendation and the cheapest seven-day test.
Name anything none of the five addressed. Do not blend the answers into
a comfortable middle if they truly conflict. Before answering, check
that your recommendation does not rely on a fact nobody supplied.
Stripping the labels mirrors the anonymised ranking step in Karpathy's repo and stops the chair deferring to a role it likes.