AI Guides › Workbench

The Four-Assistant Fit Test: Matching a Job to Claude, ChatGPT, Gemini or Grok

By Nigel Guy · 7 min read

Most people send every job to the first AI assistant they paid for, and when it fumbles, they blame "AI" rather than the mismatch. The opposite habit is just as costly: you subscribe to all four because a post told you each one "wins" at something, then can't remember which tab you were meant to open. Both habits skip the real question, which is where this particular job's inputs live and what the finished thing has to be.

The rule: choose the assistant by the job's inputs and its finished output, not by its reputation. When two tools fit equally well, run the same small sample through both before you commit.

Why the four aren't the same

All four chat, write, search, read files and write code. The real differences come from what each is plugged into and what it can hand back:

Social-media labels ("Claude for coding, Gemini for long documents") go stale fast. Anthropic's pricing page now lists a context window of "up to 1M" on every Claude plan, so long documents are no longer only Gemini's territory.

The kit at a glance

What it does best Free tier Paid entry point at time of writing Best for The catch
Claude Multi-step work across files, code and connected apps Yes: chat, web search, file creation, connectors. Sonnet and Haiku models only Pro, listed at $20 a month in US dollars. Claude Code and Research need Pro or above Drafting from your own documents, coding, jobs that take several steps No image or video generation. I couldn't find a £ price on the official page, so check at checkout
ChatGPT Fast everyday chat, voice, images Yes: voice, limited images, limited deep research Go, then Plus. No £ price showed on the official page when I checked Quick answers, voice conversations, image creation, a second opinion Free has a 27K-token context window on the Instant model, so a long upload can be only partly read
Gemini Work inside Google apps and very long inputs Yes: image generation, Deep Research, Gemini Live Google AI Plus £4.49 a month; Google AI Pro £18.99 a month Anyone whose email and documents are already in Google Workspace The full Gmail and Docs integration and the 1M context window are on Pro
Grok Live answers from the web and X Yes, with lower limits SuperGrok. I couldn't load a current official price, so check grok.com before buying Breaking news, what people are saying right now, trend checks A live X post is not a verified fact, and Grok will cite chatter as readily as reporting

How to use each one

Claude: when the job is a chain of steps

Use Claude when the job has stages: read three files, pull the figures, write the summary, build the document. Put repeated instructions in a project so you don't paste them every time. Connect only the app the job needs, such as Google Drive or Gmail. For code, Claude Code is included from the Pro plan.

Where it wins: following a long brief to the letter, holding a consistent tone across a long document, and working through a multi-file task without losing the thread.

Gemini: when the inputs already sit in Google

If the contract is in Google Docs and the thread is in Gmail, Gemini can reach both without you copying anything out. On Google AI Pro you also get the 1 million token context window and video generation with Veo 3.1 Lite (Google describes this as a limited trial).

Where it wins: summarising a long Gmail thread, drafting in Docs from material already in Drive, and very large inputs on the paid plan.

Grok: when "as of this morning" matters

Use Grok when the answer depends on the last few hours, such as reaction to a launch or a developing story. Ask it to separate primary sources from posts.

Where it wins: speed on live events and social reaction. It does not win on accuracy by default, so treat its output as a lead to check, not an answer.

ChatGPT: when you want it quick, spoken or visual

ChatGPT suits short back-and-forth work, voice conversations (on all plans, expanded on paid ones) and image creation. It also makes a good second reader for another tool's draft.

Where it wins: low-friction everyday questions, talking a problem through out loud, and images in the same chat.

The router prompt

Paste this into whichever assistant you already have open. Fill in the job, the inputs and the output you need.

You are helping me choose an AI assistant for one specific job. Be practical, not promotional.

The job: [DESCRIBE THE TASK IN ONE OR TWO SENTENCES, E.G. "SUMMARISE 40 CUSTOMER EMAILS INTO A ONE-PAGE ISSUES LIST"]
Where the inputs live: [E.G. GMAIL, LOCAL PDFS, A GITHUB REPO, LIVE NEWS]
Rough size of the inputs: [E.G. 3 PAGES, 200 PAGES, A WHOLE CODEBASE]
What I need back: [E.G. A WORD DOCUMENT, A SPOKEN ANSWER, AN IMAGE, WORKING CODE]
Tools I already pay for: [LIST, OR "NONE"]

Steps:
1. If any field above is blank or vague, ask me about it before you go on. Do not fill it in yourself.
2. Judge Claude, ChatGPT, Gemini and Grok on three things only: can it reach the inputs, can it handle their size on the plan I have, and can it produce the output format I need.
3. Name the best fit and give the deciding reason in one sentence.
4. If two or more fit equally well, say so plainly and suggest a short side-by-side test. Do not pick one just to sound decisive.
5. Flag any plan feature you are unsure is current, and tell me to check it on the vendor's pricing page.

Output: a three-row table (Reach, Size, Output) with a tick or cross per assistant, then your recommendation in no more than three sentences.

Before answering, check: did you rely on reputation rather than the three criteria? If so, redo step 2.

The side-by-side test

When the router says "either would do", spend ten minutes on this instead of reading another comparison post. Paste the same prompt into both tools, then paste both answers into a third chat with this:

You are reviewing two AI-written answers to the same task. You do not know which tool wrote which.

The task was: [PASTE THE ORIGINAL TASK]
What a good result looks like: [E.G. "ACCURATE FIGURES, UNDER 300 WORDS, PLAIN ENGLISH"]
Answer A: [PASTE]
Answer B: [PASTE]

Steps:
1. Check each answer against the success criteria, one criterion at a time.
2. List any factual claim in either answer that you cannot confirm from the task material, quoting it exactly.
3. Say which answer needs less fixing before I could use it, and why.

Output: a short table (criterion, A, B), then one paragraph with your verdict.
Do not reward length or confident tone. If both are equally usable, say so.

Keep the winner for that type of job, and don't rerun the test until the job changes.

How to choose

  1. Inputs first. Google Workspace points to Gemini, live events to Grok, local files, code or a multi-step brief to Claude, voice and images to ChatGPT.
  2. Size against your plan, not the headline figure. Free tiers sit well below the advertised context windows.
  3. Output format. Need an image or video? Claude is out.
  4. Tie? Run the side-by-side test, then keep the one you already pay for.

What to skip

Guardrails

Sources

All 751 AI guides · JulieMango plans from £17/mo