Comparison Analyze Assistant

Claude vs Grok: I ran both on a question Google gets wrong

Claude vs Grok, both run on the same two prompts. Both got the hard question right, but Grok opened 3 vendor pages while Claude cited 9 third-party posts.

Claude vs Grok: I ran both on a question Google gets wrong
Contents

Same answer, different evidence: the verdict up front

Google’s own AI Overview for Claude vs Grok says Claude wins reasoning and writing while Grok wins speed and real-time data. I put both through the same two prompts on the same afternoon, and that framing did not survive contact.

ClaudeGrok
Best forProse, long reasoning, anything with a bylineCheckable facts about a named company
Best free tierAssistant only, no model choiceFive-mode picker, including a reasoning mode
Best overall valuePro at $20, or $17 billed annuallyFree, if the free tier covers you

Both got a hard factual question right. Both obeyed all four rules on a constrained writing task. The thing that actually separated them was not accuracy and not speed. It was which sources each one decided to trust.

See Claude’s plans

How I tested this, and what that does not prove

Two prompts, run once each, in live logged-in sessions on 27 August 2026. Every screenshot below is from those runs.

One asymmetry matters and I am not going to bury it. Grok ran on a free account in its default mode, which is Fast on Grok 4.5. Claude ran on a paid Max account on Opus 5 at high effort. That is not a like-for-like capability test, and anyone telling you a free tier and a $100-a-month plan are a fair fight is selling something.

What it is like-for-like on is the question most people actually have: what does each one give me when I open it and type, without changing any settings. That is the comparison below.

The other limit is sample size. Two prompts, two tools, four answers. Run once. This is an anecdote with receipts, not a benchmark, and I would not extrapolate from it to a general claim about which model is smarter. What it can do is show you two specific behaviours in enough detail to be useful.

I also did not test SuperGrok, the paid tier. Its pricing below is first-hand from the upgrade screen, but I generated no paid Grok output and will not characterise any.

Claude vs Grok, side by side on twelve axes

AxisClaudeGrok
Free tierYes, assistant onlyYes, with a five-mode picker
Free model choiceNoneFast, Build, Auto, Expert, Heavy
Entry paid planPro, $20/moSuperGrok, $30/mo
Annual discount on entry planYes, $17/mo equivalentYes, $300/yr
Top consumer tierMax, $100 or $200/moSuperGrok is the top consumer tier
Annual billing at the topNo, Max is monthly onlyYes
Age confirmation to startNoYes
Searched the web unpromptedYesYes
Opened the vendor’s own pagesNoYes
Sources cited on the test question9 results, mostly third-party3 pages opened, 30 sources
Constraints obeyed, out of four44
Word-count accuracy on a 150-word ask152140

Two rows in that table carry most of the argument: “opened the vendor’s own pages” and “sources cited”. The rest is context.

Grok: five modes, an age gate, and no pricing page

The free tier is more generous than I expected, and the reputation for Grok being the loose, unserious one does not match what the account actually offers.

Opening the model picker on a free account gives five entries: Fast on Grok 4.5, Build on Grok 4.6 marked beta, Auto which chooses between Fast and Expert, Expert which thinks harder on Grok 4.5, and Heavy, described as a team of experts. Compare that to ChatGPT’s free tier, which our ChatGPT review documents as having no picker at all, and to Gemini’s four selectable modes covered in our ChatGPT vs Gemini comparison.

A caveat I want to be precise about. The presence of a mode in the picker is not proof of unlimited use of it. Grok’s own upgrade screen lists “smarter answers in Expert mode” as a paid benefit, which implies the free version of Expert is metered or capped in some way the interface does not spell out. Read those five entries as access, not as an allowance.

Two rough edges. Grok made me confirm my birth year before it would answer anything at all, which is a gate neither Claude nor ChatGPT imposed on me. And grok.com/pricing is a 404 — the plan prices live behind the in-app upgrade screen, so there is no public pricing page to link a colleague to.

Try Grok free

Claude: weaker at the free end, stronger above it

Claude’s free tier is the weaker of the two on paper. There is no model picker, and the free assistant is the product rather than a doorway to a range of them.

Where Claude earns its keep is above that line, and our Claude review rates it 4.4 on the strength of daily use on a paid plan. The thing it is genuinely best at in this category is prose that holds a shape over length — style guides with prohibitions in them, long drafts where the last third has to be as disciplined as the first.

The structural quirk in Claude’s pricing is at the top rather than the bottom. Pro at $20 discounts to $17 if you pay annually, but Max at $100 and $200 is monthly only, with no annual option at any price. Our Claude Max plan guide works through what those multipliers actually measure, which is a five-hour session rather than a month.

Its weakness against Grok on the test below is not intelligence. It is that when Claude searched, it did not go to the source.

Claude vs Grok pricing: which one costs less?

Grok has one consumer paid plan. Claude has three rungs. That shape difference matters more than any single price.

Screenshot of Grok's SuperGrok upgrade screen showing three hundred dollars per year with the yearly billing toggle enabled

SuperGrok is $30 a month, or $300 a year. Both figures are from the in-app upgrade screen, captured above. Twelve months at $30 would be $360, so the annual rate saves $60, a discount of about 17 percent.

Claude Pro is $20 a month, or $17 a month billed annually, which is a 15 percent discount. Claude Max is $100 or $200 a month, monthly only. Those are from Claude’s pricing page.

Read those together and the picture is not “one is cheaper”. It is that they are shaped for different buyers. Claude’s entry rung undercuts Grok’s by a third, at $20 against $30, and if all you want is a competent paid assistant that is the cheaper door. Grok’s single plan means there is no upgrade decision to agonise over, and no equivalent of the $100 question our ChatGPT Plus vs Pro guide had to untangle.

Grok has no equivalent page to link — no public pricing page exists, hence the screenshot above rather than a URL.

Grok also wins a small structural point that surprised me: its consumer plan can be paid annually, while Claude’s top tiers cannot be paid annually at all. The discount exists at the bottom of Claude’s ladder and the top of Grok’s, because Grok’s ladder only has one rung.

One more observation, offered as an observation rather than a conclusion. Grok showed me its prices in US dollars. ChatGPT’s pricing page, opened from the same machine on the same day, geo-locked to my local currency and never showed me a dollar figure at all. I am not going to convert anything or claim one approach is better. It is simply worth knowing that the price you see depends on the vendor’s choice as much as on the plan.

Infographic scorecard scoring Claude against Grok on the seven axes from my own test that could be checked objectively

The test: both got the hard question right

Here is the question I set them, and why it is a good one:

How much does ChatGPT Pro cost, and how many Pro tiers are there?

This is a trap for stale sources. OpenAI now sells Pro at two prices, $100 and $200, and a great deal of the writing on page one of Google still describes Pro as a single $200 plan. I know the correct answer cold, because I read OpenAI’s help-centre article on Pro tiers and its pricing page at source the day before and wrote them up in our ChatGPT Plus vs Pro guide.

Grok answered correctly, in 11 seconds.

Screenshot of Grok correctly identifying two ChatGPT Pro tiers at one hundred and two hundred dollars, citing OpenAI's help centre

It ran three searches and opened three pages: OpenAI’s help-centre article on Pro tiers, chatgpt.com/pricing, and the Pro plan page. Then it gave both prices, both multipliers, the monthly-only billing constraint, and the detail that the pricing page presents Pro as “From $100/month”. Every one of those matches what I verified at the source myself.

Claude also answered correctly.

Screenshot of Claude correctly identifying two ChatGPT Pro tiers, citing third-party pricing articles

Same headline figures, same multipliers, and a sharp closing caveat that OpenAI’s pricing page “presents Pro as a single plan with a usage toggle rather than two clearly separate products” — which is exactly the confusion I ran into first-hand.

So on accuracy: a tie. Both beat page one of Google.

The difference is underneath. Grok read OpenAI’s pages. Claude cited nine results, most of them third-party pricing roundups and comparison articles, with the vendor’s page among them rather than leading. Both chains reached the same place, but one of them is checkable in a single click and the other depends on how recently a handful of aggregators updated their posts.

Infographic comparing how Grok read the vendor's own pages while Claude read third-party articles about the vendor

That has a practical consequence. Claude’s answer also included details I could not verify at OpenAI, among them a specific monthly count of deep research tasks on the top tier. I am not saying those are wrong; I have no way to know, and they may well be accurate. I am saying they came from articles about OpenAI rather than from OpenAI, and the answer gives you no way to tell which parts are which.

What the wait buys, and why I am not calling it a speed test

The speed gap everyone talks about barely showed up.

Grok reported working for 11 seconds on the pricing question. Claude thought for 14 seconds on the writing task. Read those as two samples rather than a race. They come from different prompts, and I did not record the matching times for Grok on the writing task or Claude on the pricing one, so this is not a controlled speed test.

What it does show is that neither made me wait in a way that would change how I work, which is the only speed question most people actually have.

Where the two genuinely diverge is in what happens during those seconds. Grok’s trace is a browsing trace: searches run, pages opened, a line saying it was confirming pricing details. You can see the work and audit it. Claude’s is a thinking trace followed by a citation list, which tells you it reasoned but not where the reasoning touched ground.

For a factual lookup, the browsing trace is more useful, because the thing you want to check is the sources. For a writing task, the thinking time is the thing you are paying for, and there is nothing to audit.

Around the chat boxClaudeGrok
Saved workspacesProjectsProjects
Generated outputsArtifactsImagine
Runs on a scheduleScheduledAutomations
Reaches your other toolsNot a sidebar entrySkills and Connectors
Second working modeCoworkBuild

The wider workflow difference is what each product wraps around the chat box, and the two have made visibly different bets. Grok’s sidebar offers Imagine, Automations, Skills and Connectors and Projects, with a Build mode in the model picker. Claude’s offers Projects, Artifacts, Scheduled and Customize, with a Cowork mode alongside Chat.

Reaching other tools is a named sidebar entry on Grok and not on Claude, which is a statement about the menus rather than about what either can do. Grok is reaching toward acting inside other tools, and its paid pitch is an agent that signs into your Gmail and GitHub and comes back with the work done. Claude’s surfaces lean toward producing and organising artefacts you then use yourself.

Neither of those bets is settled, and I did not test either, so treat that paragraph as a description of the menus rather than a verdict on what is behind them.

One small thing the transcripts do differently. After answering, Grok offered three follow-up prompts, including one that read “Rewrite text to strictly follow all constraints” — an odd suggestion given it had followed all of them, and a hint that the interface guesses at follow-ups without checking the answer above it. Claude offered none, leaving the next move to me.

The writing test, and one strange artefact

Second prompt, four rules, all objectively checkable:

Write 150 words on why a small team might pick a cheaper AI assistant. Rules: no bullet points, no em dashes, no sentence starting with “But”, and the last word must be “budget”.

I expected this to be where Claude pulled clear, since instruction-following under constraint is the axis its reputation rests on. It did not go that way.

ClaudeGrok
No bullet pointsPassPass
No em dashesPassPass
No sentence starting with “But”PassPass
Last word is “budget”PassPass
Words delivered against 150152140
Commas used70

Both scored four out of four. Neither broke a stated rule.

Screenshot of Grok's constrained writing answer, 140 words ending on the word budget

Grok’s came in at 140 words against the 150 asked for, and landed the mandatory final word cleanly. The odd part is the punctuation: 140 words containing not one comma. Nothing in the prompt banned commas. The likeliest explanation is that it generalised the em-dash prohibition into a broader avoidance of punctuation, which is a reasonable instinct applied too widely, and the prose reads breathlessly as a result.

Screenshot of Claude's constrained writing answer, 152 words ending on the word budget

Claude hit 152 words, two over, and wrote normally punctuated prose. It was also the more concrete of the two, inventing a five-person team and a per-seat price gap to argue with where Grok stayed at the level of “limited financial resources”. Those figures are illustrative colour the prompt invited, not researched numbers, and neither tool was asked for real ones.

That concreteness is the real difference, and it is a quality judgement rather than a rule check, so weigh it accordingly. On the part that can be scored objectively, this was a tie.

Grok is the right call if…

You check facts about specific companies. This is the strongest case, and it is the one my test actually supports. If your questions are of the form “what does this product cost now” or “what did this company announce”, the tool that opens the company’s own pages is the one you want.

You want a capable free tier with real model choice. Five modes including a reasoning mode, at no cost, is a better free offer than several rivals make.

You want one plan and no upgrade maths. SuperGrok at $30 is the whole decision. There is no five-hour window to reason about and no $100 middle rung.

Skip it if your work is prose that carries your name. Nothing in my test suggested Grok is bad at writing, but the comma artefact is the kind of quiet oddity you would have to edit out every time.

Claude is the right call if…

Your output is prose people read. This remains the clearest single-axis result across every assistant this site has tested, and the constrained-writing task did nothing to weaken it even though the rule-following was a tie.

You want the cheaper paid door. $20 against $30, or $17 annually, and it buys the model most people rate highest for written work.

You need long reasoning to hold together. The gap shows up in the last third of a long piece rather than the first paragraph, which is exactly where a two-prompt test like mine cannot see it.

Skip it if the thing you actually need is current facts about a named company, and you are not going to check the sources yourself. On that job Claude reached the right answer through weaker evidence, and the next time the aggregators are stale it may not.

Final word

The received wisdom on this comparison is that Claude is the smart one and Grok is the fast one. My session did not support either half of that cleanly. Both answered a genuinely hard factual question correctly when most of the internet still answers it wrong, both followed every constraint in a writing test designed to catch sloppiness, and neither response made me wait in a way I noticed.

What did separate them was sourcing discipline. Grok went to the vendor’s own documentation and quoted it. Claude went to articles about the vendor and reached the same answer with a weaker chain behind it. For anyone using an assistant to check facts rather than to generate prose, that is the difference worth choosing on, and it is not the difference the comparison articles are arguing about.

If you want one recommendation for a typical buyer: Claude Pro at $20, because prose quality is what most people actually buy an assistant for and it is the cheaper entry. Keep a free Grok account open in another tab for the moments when you need to know what something costs today, and check what it opened before you believe it.

For the wider field at the same price, our ChatGPT vs Claude comparison covers the two most people are choosing between, and ChatGPT alternatives has the six-tool version.

See Claude’s plans

Frequently asked questions

Is Grok better than Claude?

Neither wins outright, and the honest answer depends on which failure you are more worried about.

I ran both on the same two prompts. On a factual question where most of page one is out of date, both returned the correct answer. On a writing task with four explicit constraints, both obeyed all four. Straight accuracy did not separate them.

What separated them was sourcing. Grok opened the vendor's own help centre and pricing pages and quoted them. Claude leaned on third-party articles about the vendor, the kind of pricing roundup that goes stale within weeks.

Where the primary source is authoritative and the secondary coverage is behind, that gap matters more than model quality.

One caveat that matters for reading any of this: Grok ran on a free account in its default mode and Claude on a paid Max account, so this is not a like-for-like capability test.

So: Grok for checkable facts about a specific company, Claude for reasoning and prose.

How much does Grok cost compared to Claude?

Grok has a free tier and one consumer paid plan, SuperGrok, at $30 a month or $300 a year. Claude has a free tier, Pro at $20 a month or $17 billed annually, and Max at $100 or $200 a month.

The entry paid rungs are the interesting comparison. Claude Pro at $20 undercuts SuperGrok at $30 by a third, and Claude Pro discounts to $17 if you pay annually.

Grok's annual option is the better proportional deal on its own terms. $300 a year against $30 a month is $60 saved across twelve months, which is a saving of about 17 percent. Claude Pro's annual rate saves 15 percent by the same measure.

One structural difference worth knowing: Claude's Max tiers are monthly only, with no annual option at any price, so the discount exists at the bottom of Claude's ladder and not the top.

Does Grok have a free tier?

Yes, and it is more generous than I expected before I opened it.

A free Grok account gives you a model picker with five entries: Fast on Grok 4.5, Build on Grok 4.6 in beta, Auto which chooses between Fast and Expert, Expert which thinks harder on Grok 4.5, and Heavy described as a team of experts.

That is a wider selection than several rivals hand a non-paying account. ChatGPT's free tier offers no picker at all, which our ChatGPT review documents.

Two caveats from my own session. Grok asked me to confirm my age before it would answer anything, which is a step the others do not impose. And the presence of a mode in the picker is not proof of unlimited use of it, since the upgrade screen advertises smarter answers in Expert mode as a paid benefit. Treat the five entries as access, not as an allowance.

Which is better for writing, Claude or Grok?

Claude, on the evidence of the one constrained writing task I set them, though the margin was narrower than the reputation gap suggests.

I asked both for 150 words with four rules: no bullet points, no em dashes, no sentence beginning with the word But, and a mandatory final word. Both obeyed all four. Neither broke a single stated constraint.

The difference was in what neither was asked. Claude landed on 152 words against the 150 requested. Grok came in at 140. Claude's prose carried concrete detail, inventing a five-person team and a per-seat price gap to argue with, where Grok stayed general. Those numbers are illustrative colour, not researched figures.

The oddest artefact was punctuation. Grok's 140 words contained no commas at all, which reads strangely and suggests it over-applied the em-dash ban to punctuation generally. Nothing in the prompt asked for that.

Is Grok's real-time data actually an advantage?

On the test I ran, yes, and in a more specific way than the usual framing suggests.

The standard claim is that Grok is fresher because it reaches live data from X. What I actually observed was different and more useful: Grok went to the primary source. Asked about ChatGPT Pro pricing, it opened OpenAI's help centre article and OpenAI's pricing page directly, then quoted them.

That mattered because the question was one where the secondary coverage is wrong. OpenAI now sells Pro at two prices, and much of page one of Google still describes a single plan.

Claude reached the right answer too, but via articles about OpenAI rather than OpenAI's own pages. Same conclusion, weaker chain of evidence.

The lesson generalises: for a checkable fact about a named company, prefer whichever tool reads that company's own documentation.

Can Claude access the internet like Grok?

Yes. Claude searched the web unprompted on the factual question I asked, and cited nine results, so live retrieval is not a Grok exclusive.

The difference I observed was not capability but selection. Claude's nine citations were mostly third-party pricing roundups and comparison articles, with the vendor's own pricing page among them. Grok ran three searches, opened three vendor pages, and reported drawing on thirty sources.

Both reached the correct headline figure. Claude added several details I could not verify from the vendor's pages, including a specific monthly count of deep research tasks, and attributed them to those third-party articles.

That is the practical caution. When an assistant cites an aggregator rather than the source, the answer inherits the aggregator's staleness and its errors, and you cannot tell which parts are affected without checking yourself.

Worth knowing before you weigh this: Grok was on a free account and Claude on a paid Max plan, and each prompt ran once.

Share