Learning Objectives
- Understand what Grok 4.7 changed relative to Grok 4.6, and why the published comparison is not quite like-for-like
- Read SpaceXAI's benchmark table critically β where Grok 4.7 leads and where it clearly trails Claude Fable 5.1
- Know where you can actually use Grok 4.7 today, and where Grok 4.6 is still the model you get
What Is Grok 4.7?
Grok 4.7 is the flagship model from SpaceXAI, the company formerly called xAI, released on September 21, 2026. SpaceXAI calls it its most capable model for coding and knowledge work. It is built on a new, larger base model than Grok 4.6 and trained with a longer reinforcement learning run weighted toward tasks that take many hours to finish, and the company says it is better at checking its own work and at managing long context. It was also trained to understand the Grok Bot harness natively, which SpaceXAI says makes it better at conversational and general knowledge work.
The developer documentation gives it a 500,000-token context window. Price is unchanged from Grok 4.6 at two dollars per million input tokens and six dollars per million output, so the practical argument for moving is the same one Grok 4.6 made over Grok 4.5: the capability step costs nothing extra.
Grok 4.7 is available in Cursor, in Grok Build, and through the Grok API, as well as through third-party coding harnesses, model routers and cloud platforms. It is not yet the model inside the Grok app: on the day after launch, grok.com's model picker still listed Grok 4.6 behind both its Expert and Heavy modes. So a SuperGrok subscriber and an API developer are, for now, using different models.
π―Tip
Access Grok 4.7: in the API, request the model ID grok-4.7. It is also selectable in Cursor and Grok Build. A fast variant runs at twice the output speed for twice the price.
The Benchmark Picture
SpaceXAI published a comparison table at launch. Read in full it shows a clear step over Grok 4.6, a model that beats GPT-5.6 Sol on most tests, and a model that still trails Claude Fable 5.1 on coding and terminal work.
| Benchmark | Grok 4.7 (xHigh) | Grok 4.6 (High) | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0% (high) | 65.2% | 72.7% | 70.0% |
| EEBench (electrical engineering) | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
Three readings are worth separating.
Against Grok 4.6, the gain is broad but the comparison is not quite even. Every benchmark improved, and Terminal-Bench nearly doubled from 20.3 percent to 38.0 percent. But the table runs Grok 4.7 at its xHigh reasoning effort against Grok 4.6 at High, so part of each gap is extra thinking time rather than a better model. SpaceXAI did not publish a same-effort comparison.
Against GPT-5.6 Sol, Grok 4.7 wins most rows β ahead on CursorBench, electrical engineering, AA Briefcase, Terminal-Bench and the legal benchmark, and behind on DeepSWE and HealthBench. Note that GPT-5.6 Sol is no longer OpenAI's flagship: the only comparison SpaceXAI published against the newer GPT-6 Astra is GDPval, where Grok 4.7 scored 1,695 against Astra's 1,542.
Against Claude Fable 5.1, it mostly trails. Fable leads on CursorBench, AA Briefcase and HealthBench, and by a wide margin on Terminal-Bench. On GDPval, Fable 5.1 scored 1,735 to Grok 4.7's 1,695. Grok 4.7 leads Fable on DeepSWE (at high effort), electrical engineering and the legal benchmark.
β οΈWarning
Terminal-Bench is still the clearest weakness. Grok 4.7 scores 38.0% against 57.9% for Claude Fable 5.1 β a gap of about twenty points that survived a large improvement over Grok 4.6. If your workload is terminal-driven agentic work, this is the benchmark that should decide the choice, and it does not favour Grok.
πNote
Vendor-reported numbers. The entire table comes from SpaceXAI's own launch materials, including the competitor scores, and the benchmark versions are new (CursorBench 4.0 and Terminal-Bench 4.0), so these figures cannot be compared with Grok 4.6's August launch numbers on the older versions. Treat the ordering as more reliable than the exact decimals.
Where It Genuinely Leads
Two results beat both GPT-5.6 Sol and Claude Fable 5.1, and neither is a general coding test. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6 percent against 6.7 percent for Fable 5.1 and 2.5 percent for GPT-5.6 Sol, extending the legal lead Grok 4.6 already held. On EEBench, an electrical engineering evaluation, it scored 64.0 percent against 56.4 and 39.4 percent. As with Grok 4.6, professional knowledge work is where the model is strongest relative to the field.
Safety and Cybersecurity
SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack and calls it the strongest model it has tested on refusals and jailbreak resistance. It reports that Grok 4.7 topped LatchBio's biosafety benchmark at 62.4 percent, and that on its own HackerBench test of risky cyber tasks it let only 3.3 percent of risky dual-use prompts through while rarely blocking legitimate security work. Both claims are the company's own measurements. SpaceXAI has also begun giving select cybersecurity partners invite-only access to Grok 4.7's red-team capabilities for defense research.
Pricing
- $6 per million output tokens
- Model ID grok-4.7 in the Grok API
- 500,000-token context window
- Twice the output speed
- Same model quality
- Useful for user-facing loops
- No separate SpaceXAI billing needed
- Selectable from launch day
Grok 4.6: Still in the Grok App
Grok 4.6, released on August 12, 2026, was the model this page covered until Grok 4.7 replaced it at the same price. It scored 61 on the Artificial Analysis Intelligence Index at launch, level with GPT-5.6 Sol, and its widest lead was the same legal benchmark. It remains the model behind the Grok app's Expert and Heavy modes, so anyone using Grok through a SuperGrok subscription is still on Grok 4.6 until SpaceXAI moves the app over.
Grok 4.7 vs. Other Frontier Models
| Model | Headline Strength | Where It Wins | Where It Loses |
|---|---|---|---|
| Grok 4.7 (SpaceXAI) | Professional work at low cost | Legal and electrical-engineering work; price | Terminal-style agentic tasks |
| Claude Fable 5.1 Max (Anthropic) | Strongest on coding and terminal work | CursorBench, Terminal-Bench, HealthBench | Legal-domain evaluation |
| GPT-5.6 Sol Max (OpenAI) | A prior OpenAI flagship | DeepSWE and HealthBench | Most professional-domain rows |
| Grok 4.6 (SpaceXAI) | The prior flagship, still in the Grok app | Nothing in the API at the same price | Every benchmark in the table |
Related Tools
- Grok 4.5 β the flagship before Grok 4.6
- Grok β the consumer Grok chat interface, which still runs Grok 4.6
- Grok Build β SpaceXAI's terminal coding agent, where Grok 4.7 is available from launch
- Grok Bot β SpaceXAI's always-on agent product, whose harness Grok 4.7 was trained to understand
- Cursor β the coding editor where Grok 4.7 shipped on day one
Strengths
- A clear step over Grok 4.6 β every published benchmark improved, with Terminal-Bench nearly doubling
- Leads the frontier on legal and electrical-engineering work β ahead of both GPT-5.6 Sol and Claude Fable 5.1 on the Harvey Legal Agent Benchmark and EEBench
- Unchanged pricing β two dollars input and six dollars output per million tokens, the same as Grok 4.6
- Larger context β 500,000 tokens according to the developer documentation
- Wide day-one availability β Cursor, Grok Build, the Grok API, third-party harnesses and cloud platforms
Limitations and Considerations
- Terminal-Bench is a real gap β 38.0 percent against 57.9 percent for Claude Fable 5.1
- Behind Claude Fable 5.1 on most coding benchmarks β including CursorBench, where it trails by more than five points
- Uneven comparison against Grok 4.6 β the table runs Grok 4.7 at xHigh effort and Grok 4.6 at High, so the gains overstate the like-for-like improvement
- Mostly compared against an older OpenAI model β GPT-5.6 Sol rather than GPT-6 Astra, except on GDPval
- Not in the Grok app yet β SuperGrok subscribers still get Grok 4.6
- Vendor-reported benchmarks β every figure including the competitor scores comes from SpaceXAI
- Closed model β API-only, with no downloadable weights and no self-hosting option
Best Use Cases
| Task | Why Grok 4.7 |
|---|---|
| Legal and professional-domain workflows | Its widest lead over both frontier rivals is on a legal-agent benchmark |
| Engineering analysis | It leads both rivals on EEBench, an electrical engineering evaluation |
| High-volume agentic loops | Two dollars and six dollars per million tokens is well under the frontier |
| Coding inside Cursor | Shipped there on day one, and CursorBench improved by about six points over Grok 4.6 |
When to choose alternatives:
- Terminal-driven agentic work β Claude Fable 5.1, which leads Terminal-Bench by about twenty points
- Hardest coding tasks β Claude Fable 5.1, which leads CursorBench
- Self-hosting or open weights β an open-weight model such as DeepSeek V4-Pro under MIT
Getting Started
- If you use Cursor or Grok Build, Grok 4.7 is selectable there now
- For API access, point an existing client at the SpaceXAI endpoint and set the model ID to
grok-4.7 - Choose the fast variant only for user-facing interactive loops; it costs twice as much for the same model quality
- Run your own evaluation before trusting the launch table, especially at the effort setting you will actually pay for
- If you use the Grok app, expect Grok 4.6 until SpaceXAI moves the app to Grok 4.7
Key Takeaways
- Grok 4.7 shipped on September 21, 2026 as SpaceXAI's flagship, on a new, larger base model, at the same two dollars and six dollars per million tokens as Grok 4.6
- It improves on Grok 4.6 across every published benchmark, though the comparison runs Grok 4.7 at a higher reasoning effort
- It leads both GPT-5.6 Sol and Claude Fable 5.1 on the Harvey Legal Agent Benchmark and on EEBench, but trails Fable on most coding tests
- Terminal-Bench is its clearest weakness, at 38.0 percent against 57.9 percent for Claude Fable 5.1
- It is available in the API, Cursor and Grok Build; the Grok app still runs Grok 4.6
- Every figure in the comparison table is SpaceXAI's own, including the competitor scores











