Get The Signal
← Back to radar
Comparison

ChatGPT vs Claude vs Gemini: We Ran 40 Real Tasks Through All Three

Same prompts, same rubric, same grader. The honest scoreboard for the big three AI assistants in 2026 — and which one actually deserves your $20.

Everyone has an opinion on this fight. Most of it is tribal. So we did the annoying thing: we ran 40 identical real-world tasks through ChatGPT, Claude, and Gemini — the same prompts, the same scoring rubric, the same grader — and let the numbers settle it.

Spoiler: there's no single winner. There IS a clear winner for your use case, and by the end of this you'll know which one to pay for.

The test: 40 tasks, zero mercy

We split the tasks into five buckets: writing (8), coding (8), analysis & research (8), multimodal (8), and everyday assistant work (8). Each response was scored 1-10 on accuracy, usefulness, and how much editing it needed. I did the grading blind — responses were stripped of which model produced them.

Here's the thing nobody tells you: at the top end, these models are close. The differences that actually matter are the edges — context length, refusal behavior, and whether the thing remembers what you said 40 messages ago.

The scoreboard

CategoryChatGPTClaudeGemini
Writing8.49.1 ★8.0
Coding9.0 ★8.98.2
Analysis & research8.58.79.0 ★
Multimodal8.68.39.2 ★
Everyday assistant8.8 ★8.58.4
Overall8.78.78.6

Look at that overall row. They're tied. Anyone who tells you one is "obviously the best" is selling you something. The category rows are where your decision lives.

ChatGPT: the reliable generalist

GPTChatGPT8.7

ChatGPT is the Honda Civic of AI — not the flashiest, but it starts every morning and goes anywhere. Its coding edge came from better tool use and fewer 'I can't do that' refusals. The ecosystem (custom GPTs, memory, voice) is genuinely the most polished.

Where it lost points: prose that defaults to a recognizable 'AI cadence,' and a tendency to hedge when you wanted a straight answer.

What we loved

  • Best coding + tool use in the test
  • Most polished ecosystem (voice, memory, GPTs)
  • Fewest unnecessary refusals

What hurt

  • Prose has a detectable AI cadence
  • Hedges when you want a direct call
  • Free tier is more limited than rivals
Pricing: $20/mo Plus · free tier
CLDClaude8.7

Claude writes like someone who reads books. It won the writing category by the widest margin of any single category in the test, and its long-context handling was flawless — we fed it a 90-page contract and it cited the right clause every time.

Its weakness: it's more cautious. It refused two tasks the others attempted, and its multimodal work trailed Gemini noticeably.

What we loved

  • Best writing quality, full stop
  • Flawless long-context recall
  • Most natural conversational tone

What hurt

  • More refusals than competitors
  • Weaker image/video understanding
  • No native web-search depth vs Gemini
Pricing: $20/mo Pro · generous free tier
GEMGemini8.6

Gemini is the sleeper pick. It won analysis and multimodal outright — its Google Search grounding means its research answers came with sources that actually checked out, and its video understanding is in a different league.

Where it stumbled: writing felt stiff, and it occasionally mixed up details across long conversations.

What we loved

  • Best research + source grounding
  • Best-in-class multimodal (video especially)
  • Deep Google Workspace integration

What hurt

  • Stiff, corporate prose
  • Occasional context mix-ups long-conversation
  • Availability varies by region
Pricing: $20/mo Advanced · free tier

So which one do you pay for?

Writers and anyone who ships prose: Claude. It's not close.

Developers and power users: ChatGPT. The tool use and ecosystem win.

Researchers, analysts, and Google-shop people: Gemini. The grounding is real.

Broke? Use all three free tiers and rotate. In 2026 that's a genuinely viable strategy, not a joke.

FAQ

Is the $20/mo version worth it over free?
If you use it daily, yes — the free tiers throttle you right when you need them most. If it's twice a week, free is fine.
Which one hallucinates the least?
In our test, Claude produced the fewest confident false claims, followed closely by Gemini (which leans on search grounding). ChatGPT hallucinated most often but was easiest to fact-check via its sources feature.
Can I use all three for $20?
No — it's $20 each. But the free tiers plus one paid subscription covers most people's actual needs.
MK

Mara Kessler · Senior Editor

Former newsroom editor, full-time AI tool breaker. I buy every subscription myself so you don't have to trust a sponsored ranking. Reach me via the contact page — I read everything.

Frequently asked questions

How many tasks did you run across the three?

40 real tasks — writing, code, research, spreadsheets — graded blind across all three models.

Which model is best at coding?

Claude led on multi-file reasoning, ChatGPT on quick one-shot scripts, Gemini on reading huge documents.

Should I pay for all three at once?

Rarely. Pick the one matching your main task; our matrix shows exactly where each pays off.