AI Guides › Playbooks
By Nigel Guy · 7 min read
Most people who want an AI workflow start by building one: a scheduled agent, a chain of steps, a loop that runs overnight. Then it produces confident, bland, slightly wrong output on a timer. It feels like progress because something is running, but nobody checked whether the model could do the task in a plain conversation first. Automation does not fix a task you cannot steer by hand. It repeats the problem faster.
The rule: if you cannot get a good result from a chat by steering it yourself, you are not ready to automate it. Use five short probes to find out, and write down what worked.
Each probe is a single instruction you can type after the model's first answer. Each one tests a different thing, and each answer tells you what to put into the eventual automated prompt.
| Probe | What it tests | What you keep |
|---|---|---|
| 1. Range | Whether the first answer was the only option | The options worth choosing between |
| 2. Pressure | Whether the answer survives attack | The weak points and caveats |
| 3. Pace | Whether the model can work at your speed and depth | The length and tone you actually want |
| 4. Push back | Whether it will disagree instead of flatter | Evidence of honest critique |
| 5. Ask first | Whether it knows what it is missing | The inputs the task really needs |
Why this works: the prompting guidance from Anthropic and OpenAI agrees on the basic mechanism. The model cannot read your mind, so the less it has to guess, the better the result. Probes are a cheap way to find out what it was guessing at.
A first answer is one sample, not the answer. Asking for range exposes how much of it was chance.
Give me [NUMBER_OF_OPTIONS] genuinely different approaches to [TASK_OR_QUESTION]. Make them differ in strategy, not wording. For each, say in one line who it suits and what it costs me. Finish by saying which you would pick for [MY_SITUATION] and why.
Fill in: the number of options (three to five), the task, and a sentence about your situation. If the options are the same idea in different clothes, the model has no real range on this task, and that is worth knowing before you automate it.
Take the option you liked and attack it before anything else does.
You are a sceptical reviewer who has seen [TYPE_OF_PROJECT] fail before. Take the plan below and find the [NUMBER_OF_RISKS] most likely ways it breaks. For each: what goes wrong, how I would notice, and the cheapest fix. Do not soften anything. If you cannot find a real weakness, say so rather than inventing one.
Plan: [PASTE_THE_PLAN]
Fill in: the type of project, how many risks, and the plan. The last sentence of the prompt matters. Without permission to say "no real weakness", models tend to produce a list because you asked for one.
This is about tempo and depth: how fast and how dense you want the answer. It is the probe people skip, and it is the one that shapes your final output format.
Redo your last answer in [LENGTH, e.g. five bullet points / 150 words / one paragraph] for a reader who is [READER_DESCRIPTION]. Keep the one idea that matters most and drop the rest. Then tell me, in one line, what you cut and whether it mattered.
Fill in: the length and who will read it. The "what you cut" line is a self-check: it shows you whether compression lost something important. Run it twice at different lengths. The version you prefer becomes the output specification in your automation.
Models are inclined to agree with the person they are talking to. Test that directly.
I am leaning towards [MY_DECISION_OR_VIEW]. Do not agree with me by default. Give the strongest case against it, then tell me which of my assumptions you would check first. If after that you still think I am right, say so and explain what would change your mind.
My reasoning: [MY_REASONING]
Fill in: your view and the reasoning behind it. Try it once with a view you know is weak. If the model still approves, you have learned that its "looks good" is not evidence, and any automated review step built on it will be rubber-stamping.
The best test of whether a task is ready is whether the model can tell you what it lacks.
Before you answer, read the task below and list the questions you would need answered to do it well. Group them as: must-have, nice-to-have, and things you would otherwise assume. Ask me the must-haves one at a time and wait for my reply. Do not start the work until I say go.
Task: [TASK_DESCRIPTION]
Fill in: the task. The "things you would otherwise assume" group is the most useful part. Each item is a hidden guess that will bite you when nobody is watching the run. Turn every answer into a line in your final prompt.
Run them in this order, in one chat, so each builds on the last: Ask first, Range, Pressure, Push back, Pace. Ask first fills the gaps, Range gives options, Pressure and Push back stress the pick, Pace fixes the format.
A hypothetical example. You run a small bakery and want an AI to draft a weekly customer email.
You now have a prompt with real inputs, a chosen approach, a skip rule and a length. That is the thing worth automating, and you got there without building anything.
Treating the probes as a ritual. If you run all five, accept every answer, and move on, you have only added theatre. The probes pay off when you change something because of them: a rule added, an option dropped, a length fixed. The second trap is stopping at fluent. Smooth, confident prose proves nothing about whether the content is right, so check facts yourself on anything that matters.
It does not make the model correct, it does not replace testing an automation on real examples, and it does not tell you whether the cost or risk of automating is worth it. It tests steerability, not truth. Where the stakes are legal, medical or financial, a human with the right expertise still decides.