Comparison

Grok 4.7 vs Claude: task fit on xAI’s own sheet

Not a trophy. Grok 4.7 and Claude both generate. This page uses xAI’s 21 Sep 2026 table to separate electrical engineering, legal-agent work, long coding, terminal work, and clinical reasoning — and to stop a Max SKU being compared with a workhorse price.

TaskGrok 4.7Claude Fable 5.1 Max
List price, $ per 1M in / out$2 / $6$10 / $50
Longer coding (CursorBench 4.0)46.3%51.8%
Software engineering (DeepSWE v1.1)71.0%, high effort70.0%
Electrical engineering (EEBench)64.0%56.4%
Multi-hour office (AA Briefcase v1.1)1,6571,678
Terminal work (Terminal-Bench 4.0)38.0%57.9%
Legal agent (Harvey)19.6%6.7%
Clinical reasoning (HealthBench Professional)56.7%62.1%
Typed decision onlyWrong defaultWrong default

The actual question

Not “which lab is ahead this week?” Both models generate code and prose and can sit inside an agent harness. Ask: which failure is expensive, and which SKU is actually in the quote? Grok 4.7 is the cheaper list price on xAI’s sheet. Claude’s Fable 5.1 Max is ahead on several of the long-running coding and office tasks in that same sheet. Those can both be true.

Do not mix the SKUs

The Claude column in xAI’s launch table is Fable 5.1 Max at $10 input / $50 output per million tokens. That is not Claude Sonnet 5. Our Claude profile still dates Sonnet 5 at $2 / $10 and Opus 5 at $5 / $25. Beating a $50 output SKU on price is easier than beating the workhorse your team actually has pinned. If your Claude bill is Sonnet, this page does not say Grok is half the price of your invoice.

The OpenAI column next to them is GPT-5.6 Sol Max at $4 / $20. It is context, not the subject of this page. Full Grok positioning is on the Grok 4.7 profile.

Same job, split by the vendor’s tasks

Every percentage below is from xAI’s table, published with the model on 21 September 2026. We did not rerun the harnesses. Where Grok 4.7 is marked high effort, that asterisk stays.

Lean Grok 4.7 when

  • The work looks like electrical engineering as EEBench defines it. 64.0% versus Fable 5.1 Max at 56.4% is the cleanest gap in Grok’s favor on this sheet.
  • You are running a legal-agent harness in the Harvey sense and you have a human on the result. 19.6% versus 6.7% is a large relative gap and a small absolute score. Do not read it as “Grok can practice law.”
  • The constraint is list price for a Max-class comparison. $2 / $6 against $10 / $50 is the number xAI wants quoted. It is real as a list price. It is not a measured invoice.
  • You are already in Cursor or Grok Build and the migration cost is the point. Availability is part of task fit.

Lean Claude when

  • The job is terminal-heavy agent work. Terminal-Bench 4.0 is 57.9% for Fable 5.1 Max and 38.0% for Grok 4.7. That is the result that should stop a blanket “Grok for agents” switch.
  • The job is longer coding as CursorBench 4.0 scores it. Fable 5.1 Max is 51.8%. Grok 4.7 is 46.3%. Grok’s own price-performance sentence can still be true. The raw score is not a win.
  • The deliverable is multi-hour office work in the AA Briefcase sense. 1,678 versus 1,657 is close. Close is not a reason to migrate a working Claude pipeline.
  • The domain is clinical reasoning. HealthBench Professional has Fable at 62.1% and Grok 4.7 at 56.7%. Grok is not the specialist on this row.
  • You want Anthropic’s careful-reader default on a long document a person will sign. That is a product stance, discussed on the Claude profile, not a cell in this table.

Where they are close enough to ignore the leaderboard

DeepSWE v1.1 is 71.0% for Grok 4.7 at high effort and 70.0% for Fable 5.1 Max. A one-point vendor gap with an effort asterisk is not a procurement decision. AA Briefcase is the same shape of result: a small gap, a large marketing sentence available to either side. If your traces do not look like these benchmarks, neither number transfers. Measure the failure you actually pay for.

Price is a task, not a tie-break you skip

Grok 4.7 did not get cheaper than Grok 4.6. The upgrade case is capability at a flat list price, plus a fast tier at twice the output speed and twice the price. Claude’s Max SKU in the comparison is much more expensive per token and, on three of these rows, ahead. Paying $50 output for Terminal-Bench and CursorBench can be rational. Paying it because a homepage said “frontier” is not. The arithmetic, including the fast tier and the Sonnet mix-up, is on the Grok pricing review.

What we will not declare

A winner. The sheet is one lab’s exhibit, published the day the model shipped. Public argument about GPT-6 Astra and GDPval is already louder than the cells xAI printed in the main table. We stay with the cells. If you needed a typed decision instead of either generator, leave this page for Jev vs LLM.

Get listed

Put your AI tool in front of people who are already comparing options.

Submit a listing for review. Complete submissions with a live website, pricing, and a clear use case typically go live within 24–72 hours.

Submit a tool