AI Guides › Playbooks

The Five-Probe Conversation: Test a Task by Talking Before You Automate It

By Nigel Guy · 7 min read

Most people who want an AI workflow start by building one: a scheduled agent, a chain of steps, a loop that runs overnight. Then it produces confident, bland, slightly wrong output on a timer. It feels like progress because something is running, but nobody checked whether the model could do the task in a plain conversation first. Automation does not fix a task you cannot steer by hand. It repeats the problem faster.

The rule: if you cannot get a good result from a chat by steering it yourself, you are not ready to automate it. Use five short probes to find out, and write down what worked.

The Five-Probe Card

Each probe is a single instruction you can type after the model's first answer. Each one tests a different thing, and each answer tells you what to put into the eventual automated prompt.

Probe What it tests What you keep
1. Range Whether the first answer was the only option The options worth choosing between
2. Pressure Whether the answer survives attack The weak points and caveats
3. Pace Whether the model can work at your speed and depth The length and tone you actually want
4. Push back Whether it will disagree instead of flatter Evidence of honest critique
5. Ask first Whether it knows what it is missing The inputs the task really needs

Why this works: the prompting guidance from Anthropic and OpenAI agrees on the basic mechanism. The model cannot read your mind, so the less it has to guess, the better the result. Probes are a cheap way to find out what it was guessing at.

Probe 1: Range

A first answer is one sample, not the answer. Asking for range exposes how much of it was chance.

Give me [NUMBER_OF_OPTIONS] genuinely different approaches to [TASK_OR_QUESTION]. Make them differ in strategy, not wording. For each, say in one line who it suits and what it costs me. Finish by saying which you would pick for [MY_SITUATION] and why.

Fill in: the number of options (three to five), the task, and a sentence about your situation. If the options are the same idea in different clothes, the model has no real range on this task, and that is worth knowing before you automate it.

Probe 2: Pressure

Take the option you liked and attack it before anything else does.

You are a sceptical reviewer who has seen [TYPE_OF_PROJECT] fail before. Take the plan below and find the [NUMBER_OF_RISKS] most likely ways it breaks. For each: what goes wrong, how I would notice, and the cheapest fix. Do not soften anything. If you cannot find a real weakness, say so rather than inventing one.

Plan: [PASTE_THE_PLAN]

Fill in: the type of project, how many risks, and the plan. The last sentence of the prompt matters. Without permission to say "no real weakness", models tend to produce a list because you asked for one.

Probe 3: Pace

This is about tempo and depth: how fast and how dense you want the answer. It is the probe people skip, and it is the one that shapes your final output format.

Redo your last answer in [LENGTH, e.g. five bullet points / 150 words / one paragraph] for a reader who is [READER_DESCRIPTION]. Keep the one idea that matters most and drop the rest. Then tell me, in one line, what you cut and whether it mattered.

Fill in: the length and who will read it. The "what you cut" line is a self-check: it shows you whether compression lost something important. Run it twice at different lengths. The version you prefer becomes the output specification in your automation.

Probe 4: Push back

Models are inclined to agree with the person they are talking to. Test that directly.

I am leaning towards [MY_DECISION_OR_VIEW]. Do not agree with me by default. Give the strongest case against it, then tell me which of my assumptions you would check first. If after that you still think I am right, say so and explain what would change your mind.

My reasoning: [MY_REASONING]

Fill in: your view and the reasoning behind it. Try it once with a view you know is weak. If the model still approves, you have learned that its "looks good" is not evidence, and any automated review step built on it will be rubber-stamping.

Probe 5: Ask first

The best test of whether a task is ready is whether the model can tell you what it lacks.

Before you answer, read the task below and list the questions you would need answered to do it well. Group them as: must-have, nice-to-have, and things you would otherwise assume. Ask me the must-haves one at a time and wait for my reply. Do not start the work until I say go.

Task: [TASK_DESCRIPTION]

Fill in: the task. The "things you would otherwise assume" group is the most useful part. Each item is a hidden guess that will bite you when nobody is watching the run. Turn every answer into a line in your final prompt.

Using all five in one conversation

Run them in this order, in one chat, so each builds on the last: Ask first, Range, Pressure, Push back, Pace. Ask first fills the gaps, Range gives options, Pressure and Push back stress the pick, Pace fixes the format.

A hypothetical example. You run a small bakery and want an AI to draft a weekly customer email.

  1. Ask first: the model asks about your tone, whether you want offers, and what the last email said. You discover you have never defined your tone.
  2. Range: it offers a menu-led email, a story-led email and a "one thing this week" email. You pick the third.
  3. Pressure: it warns that "one thing" emails go stale if you have nothing new, and that discounts every week train customers to wait. You add a rule: skip the week if nothing is new.
  4. Push back: you say you want to include a 20% offer every time. It argues against, you disagree, and it notes what would change its mind. Honest friction is the sign you want.
  5. Pace: you ask for 120 words for a regular customer reading on a phone, then compare against 200.

You now have a prompt with real inputs, a chosen approach, a skip rule and a length. That is the thing worth automating, and you got there without building anything.

What's the trap?

Treating the probes as a ritual. If you run all five, accept every answer, and move on, you have only added theatre. The probes pay off when you change something because of them: a rule added, an option dropped, a length fixed. The second trap is stopping at fluent. Smooth, confident prose proves nothing about whether the content is right, so check facts yourself on anything that matters.

What won't this do?

It does not make the model correct, it does not replace testing an automation on real examples, and it does not tell you whether the cost or risk of automating is worth it. It tests steerability, not truth. Where the stakes are legal, medical or financial, a human with the right expertise still decides.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo