Grok 4.7: a coding and knowledge-work model, not a trophy
xAI’s Grok 4.7 is a general-purpose generator aimed at long coding tasks and professional knowledge work. This profile dates the vendor table, labels the slogan, and says which jobs the sheet actually supports.
What Grok 4.7 is
Grok 4.7 is xAI’s current general-purpose model for coding and knowledge work. It generates text, code, and tool-using runs. It is not a decision model. If the only output you need is a closed label, route, or allow/deny, this is the wrong default — read Jev before you pay generation prices for a choice.
xAI’s launch line is “twice as fast, at half the price of comparable models,” and also “the same price and speed as Grok 4.6.” Those sentences are not the same claim. The second one is the useful one: the list price did not drop versus 4.6. The first one depends on which model you call comparable. On xAI’s own sheet, Fable 5.1 Max is $10 / $50 and GPT-5.6 Sol Max is $4 / $20, against Grok 4.7 at $2 / $6. Half of Fable’s list price is a real gap. Half of every frontier model is a slogan. We label it as a vendor claim.
Developer
xAI. The model is called through the Grok API, inside Cursor and Grok Build, and through third-party coding harnesses, routers, and cloud platforms that xAI listed on launch day. GitHub Copilot started a gradual rollout the same day, billed at provider list price. Consumer Grok chat is a different surface from the API bill.
What changed versus Grok 4.6
xAI says 4.7 uses a new, larger base model, trained with a longer reinforcement-learning run on a harder mix, weighted toward problems that take many hours. The vendor also says it is better at checking its own work, handling longer context, and driving the Grok Bot harness for conversational and general knowledge work. Documents and presentations are called out as improved. None of that is an AIToolsRatings measurement. It is the reason the vendor gives for shipping a new id at the old price.
We are not printing parameter counts or context-window sizes. They are not in the launch note we are citing. Secondary writeups disagree with each other. A profile that invents a context length is worse than one that waits.
Where the vendor sheet is strong
Read the table below as xAI’s comparison, dated 21 September 2026, not as our rerun. On that sheet Grok 4.7 beats Grok 4.6 on every listed task. Against the other two columns it does not sweep:
- Electrical engineering (EEBench): 64.0%, ahead of Fable 5.1 Max at 56.4% and GPT-5.6 Sol Max at 39.4%. This is the clearest vendor win on the page.
- Legal agent work (Harvey): 19.6%, ahead of Fable at 6.7% and Sol at 2.5%. The absolute number is still low. “Best on a hard legal-agent bench” is not “ready to file.”
- DeepSWE v1.1: 71.0% at high effort, a hair above Fable’s 70.0% and below Sol’s 72.7%. The asterisk is theirs. Do not drop it when you quote the number.
- CursorBench 4.0: 46.3%. Better than 4.6 (40.4%) and Sol Max (41.7%). Behind Fable 5.1 Max at 51.8%. xAI still calls the price-performance “at the frontier” here. Price-performance is the claim, not the raw score.
Where the same sheet is behind
- Terminal-Bench 4.0: 38.0% versus Fable 5.1 Max at 57.9%. Sol Max is 37.3%, so this is not a Grok-versus-everyone loss, but it is a large gap to the Claude SKU in the table.
- AA Briefcase v1.1: 1,657 versus Fable at 1,678. Close, and ahead of Sol at 1,487. “Comparable” is a fair word. “Wins professional knowledge work” is not.
- HealthBench Professional: 56.7%, behind both Sol (60.5%) and Fable (62.1%). Clinical reasoning is not the reason to switch.
xAI also points at GDPval for professional knowledge work and says 4.7 improves on 4.6 and lands near other frontier models. The launch note’s chart names GPT-6 Astra on that view. We are not copying a single GDPval integer into this page until it is as explicit as the table above. The chart is a vendor exhibit. Treat it that way.
Safety claims
xAI says 4.7 has a new safeguard stack and is the strongest model it has tested on refusals and jailbreak resistance. Two numbers are specific enough to date: 62.4% on LatchBio’s biosafety benchmark, which the vendor says it leads, and 3.3% of risky dual-use prompts allowed through on HackerBench v0.3, with a claim that legitimate security work is rarely blocked. Invite-only red-team access for selected cybersecurity partners is a distribution choice, not a public feature. We have not reproduced any of these tests.
Input and output
- Input and output are billed per million tokens. The published start is $2 in / $6 out.
- A fast variant is served at twice the output speed and, in xAI’s words, twice the price. Read against $2 / $6, that is $4 / $12 until a rate card says otherwise.
- Longer context and self-checks are vendor claims. They can also mean more output tokens. We have not metered that.
- Structured tool calls do not turn the model into a classifier.
When to pick something else
- The artifact is a long document a human will ship, and you want Claude’s careful-reader default → Claude, then the task-fit comparison.
- The audience already lives in ChatGPT, or you need that plugin graph → ChatGPT.
- You only need a typed decision at high QPS → Jev.
- You are choosing on list price alone → open the pricing review before you treat $2 / $6 as the whole bill.
Latest update
22 September 2026 desk note on the 21 September launch. Grok 4.7 is live in Cursor, Grok Build, and the API, with Copilot still rolling out. This page is a profile. It is not a setup guide, and it is not a leaderboard crown.
Task fit, not a winner
| Task (xAI table, 21 Sep 2026) | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| Input / output, $ per 1M tokens | $2 / $6 | $2 / $6 | $4 / $20 | $10 / $50 |
| CursorBench 4.0, longer coding | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0% (high effort) | 65.2% | 72.7% | 70.0% |
| EEBench, electrical engineering | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1, multi-hour office | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |