AI Guides › Workbench
By Nigel Guy · 7 min read
Most people send every job to the first AI assistant they paid for, and when it fumbles, they blame "AI" rather than the mismatch. The opposite habit is just as costly: you subscribe to all four because a post told you each one "wins" at something, then can't remember which tab you were meant to open. Both habits skip the real question, which is where this particular job's inputs live and what the finished thing has to be.
The rule: choose the assistant by the job's inputs and its finished output, not by its reputation. When two tools fit equally well, run the same small sample through both before you commit.
All four chat, write, search, read files and write code. The real differences come from what each is plugged into and what it can hand back:
Social-media labels ("Claude for coding, Gemini for long documents") go stale fast. Anthropic's pricing page now lists a context window of "up to 1M" on every Claude plan, so long documents are no longer only Gemini's territory.
| What it does best | Free tier | Paid entry point at time of writing | Best for | The catch | |
|---|---|---|---|---|---|
| Claude | Multi-step work across files, code and connected apps | Yes: chat, web search, file creation, connectors. Sonnet and Haiku models only | Pro, listed at $20 a month in US dollars. Claude Code and Research need Pro or above | Drafting from your own documents, coding, jobs that take several steps | No image or video generation. I couldn't find a £ price on the official page, so check at checkout |
| ChatGPT | Fast everyday chat, voice, images | Yes: voice, limited images, limited deep research | Go, then Plus. No £ price showed on the official page when I checked | Quick answers, voice conversations, image creation, a second opinion | Free has a 27K-token context window on the Instant model, so a long upload can be only partly read |
| Gemini | Work inside Google apps and very long inputs | Yes: image generation, Deep Research, Gemini Live | Google AI Plus £4.49 a month; Google AI Pro £18.99 a month | Anyone whose email and documents are already in Google Workspace | The full Gmail and Docs integration and the 1M context window are on Pro |
| Grok | Live answers from the web and X | Yes, with lower limits | SuperGrok. I couldn't load a current official price, so check grok.com before buying | Breaking news, what people are saying right now, trend checks | A live X post is not a verified fact, and Grok will cite chatter as readily as reporting |
Use Claude when the job has stages: read three files, pull the figures, write the summary, build the document. Put repeated instructions in a project so you don't paste them every time. Connect only the app the job needs, such as Google Drive or Gmail. For code, Claude Code is included from the Pro plan.
Where it wins: following a long brief to the letter, holding a consistent tone across a long document, and working through a multi-file task without losing the thread.
If the contract is in Google Docs and the thread is in Gmail, Gemini can reach both without you copying anything out. On Google AI Pro you also get the 1 million token context window and video generation with Veo 3.1 Lite (Google describes this as a limited trial).
Where it wins: summarising a long Gmail thread, drafting in Docs from material already in Drive, and very large inputs on the paid plan.
Use Grok when the answer depends on the last few hours, such as reaction to a launch or a developing story. Ask it to separate primary sources from posts.
Where it wins: speed on live events and social reaction. It does not win on accuracy by default, so treat its output as a lead to check, not an answer.
ChatGPT suits short back-and-forth work, voice conversations (on all plans, expanded on paid ones) and image creation. It also makes a good second reader for another tool's draft.
Where it wins: low-friction everyday questions, talking a problem through out loud, and images in the same chat.
Paste this into whichever assistant you already have open. Fill in the job, the inputs and the output you need.
You are helping me choose an AI assistant for one specific job. Be practical, not promotional.
The job: [DESCRIBE THE TASK IN ONE OR TWO SENTENCES, E.G. "SUMMARISE 40 CUSTOMER EMAILS INTO A ONE-PAGE ISSUES LIST"]
Where the inputs live: [E.G. GMAIL, LOCAL PDFS, A GITHUB REPO, LIVE NEWS]
Rough size of the inputs: [E.G. 3 PAGES, 200 PAGES, A WHOLE CODEBASE]
What I need back: [E.G. A WORD DOCUMENT, A SPOKEN ANSWER, AN IMAGE, WORKING CODE]
Tools I already pay for: [LIST, OR "NONE"]
Steps:
1. If any field above is blank or vague, ask me about it before you go on. Do not fill it in yourself.
2. Judge Claude, ChatGPT, Gemini and Grok on three things only: can it reach the inputs, can it handle their size on the plan I have, and can it produce the output format I need.
3. Name the best fit and give the deciding reason in one sentence.
4. If two or more fit equally well, say so plainly and suggest a short side-by-side test. Do not pick one just to sound decisive.
5. Flag any plan feature you are unsure is current, and tell me to check it on the vendor's pricing page.
Output: a three-row table (Reach, Size, Output) with a tick or cross per assistant, then your recommendation in no more than three sentences.
Before answering, check: did you rely on reputation rather than the three criteria? If so, redo step 2.
When the router says "either would do", spend ten minutes on this instead of reading another comparison post. Paste the same prompt into both tools, then paste both answers into a third chat with this:
You are reviewing two AI-written answers to the same task. You do not know which tool wrote which.
The task was: [PASTE THE ORIGINAL TASK]
What a good result looks like: [E.G. "ACCURATE FIGURES, UNDER 300 WORDS, PLAIN ENGLISH"]
Answer A: [PASTE]
Answer B: [PASTE]
Steps:
1. Check each answer against the success criteria, one criterion at a time.
2. List any factual claim in either answer that you cannot confirm from the task material, quoting it exactly.
3. Say which answer needs less fixing before I could use it, and why.
Output: a short table (criterion, A, B), then one paragraph with your verdict.
Do not reward length or confident tone. If both are equally usable, say so.
Keep the winner for that type of job, and don't rerun the test until the job changes.