Twice the Work, Same Price Tag: Inside SpaceXAI's Grok 4.7
SpaceXAI's Grok 4.7 runs longer on hard tasks, nearly doubles Terminal-Bench scores, and posts 19.6% on Harvey's legal agent benchmark — all at Grok 4.6's $2/$6 pricing, though its gains come with more than double the token use.
On September 21, 2026, SpaceXAI quietly shipped Grok 4.7, its most capable model yet for coding, agentic tasks, and knowledge work — and it did so without touching the price list. Where rivals spent the same week cutting prices or racing each other to launch (OpenAI shipped GPT-6 Sol and Luna roughly ninety minutes after Anthropic’s Opus 5.5), SpaceXAI’s move was the opposite: keep the $2 per million input / $6 per million output pricing of Grok 4.6, and make the model behind it substantially better. “Twice as fast, at half the price of comparable models,” is the company’s framing, and the benchmark table it published backs at least the second half of that claim.
What changed under the hood
Four things separate Grok 4.7 from its predecessor, and none of them are incremental:
A new, larger base model. Grok 4.7 does not reuse the Grok 4.6 base. SpaceXAI trained a bigger foundation model for this generation — a notable decision in a year when most labs have squeezed gains out of existing bases through post-training alone.
A longer RL run on harder tasks. The reinforcement learning mix was deliberately weighted toward problems that take many hours to complete. That choice shows up directly in the long-horizon benchmark results: this is a model tuned for the reality of agentic work, where a single task spans hours rather than seconds.
Better self-verification and long-context handling. The model checks its own work more carefully — a claim that matters most in coding, where catching your own mistakes mid-run is the difference between a working PR and a hallucinated one.
Native Grok Bot harness support. Grok 4.7 was trained to natively understand SpaceXAI’s Grok Bot harness, improving conversational quality and general knowledge work.
The developer docs round out the specs: a 500,000-token context window (unchanged from Grok 4.6), knowledge cutoff of May 2026, text and image input with text output, and a configurable reasoning effort spanning low to xhigh.
The benchmark table
SpaceXAI’s launch comparison pits Grok 4.7 at xHigh effort against Grok 4.6 High, GPT-5.6 Sol Max, and Anthropic’s Fable 5.1 Max. All scores are vendor-reported, and the honest summary is: Grok 4.7 improves on Grok 4.6 in every single row, leads on two benchmarks outright, but does not sweep the board.
| Benchmark | Grok 4.7 xHigh | Grok 4.6 High | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| Input price ($/M) | $2 | $2 | $4 | $10 |
| Output price ($/M) | $6 | $6 | $20 | $50 |
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 37.6% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
*High effort.
Three numbers stand out. The Terminal-Bench 4.0 jump from 20.3% to 37.6% (Artificial Analysis independently measured a +4.5 point gain on its own run) is the single largest generational improvement in the table — long-running terminal work was Grok 4.6’s weakest suit and is now at parity with GPT-5.6 Sol. The EEBench score of 64.0% tops the table by nearly 8 points over Fable 5.1, and dwarfs GPT-5.6 Sol’s 39.4% — electrical engineering is an odd place to find a frontier lead, and SpaceXAI is not shy about it. And the Harvey Legal Agent Benchmark result of 19.6% is close to a category win: Fable 5.1 Max manages 6.7% and GPT-5.6 Sol Max just 2.5%, meaning Grok 4.7 is roughly triple its nearest competitor’s legal-agent score at a third of the price.
But Fable 5.1 Max still tops 4 of 7 benchmarks, including a dominant 57.9% on Terminal-Bench — and it costs 5x more on input and ~8.3x more on output. GPT-5.6 Sol Max holds the best DeepSWE result at 72.7%. On SpaceXAI’s cost-per-task chart for CursorBench 4.0, Grok 4.7 sits at the frontier of price-performance. That is the actual pitch: not “best model,” but “frontier capability at a mid-tier price.”
Independent verification from Artificial Analysis
Vendor benchmarks deserve skepticism, which makes Artificial Analysis’s independent evaluation the most useful data point in the release. Their verdict is broadly confirmatory, with one important caveat.
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index, up 2 points over Grok 4.6, which AA says brings SpaceXAI “into the top 4 AI labs” by its ranking. On AA-Briefcase, AA’s private benchmark for long-horizon agentic knowledge work, Grok 4.7 gains +111 Elo over its predecessor to reach 1,657 — placing it just behind Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1,695 Elo, +90 ahead of Grok 4.6.
The coding story is stronger: Grok 4.7 + Grok Build scores 56 on AA’s Coding Agent Index, up 9 points from Grok 4.6, ranking 4th among models in their native harnesses — behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. Component scores improved across the board: DeepSWE v1.1 from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, SWE-Atlas-QnA from 58% to 63%.
The caveat: the gains come with a big token appetite. Grok 4.7 (xhigh) burns roughly 81k output tokens per Intelligence Index task — versus 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max). That is 125% more than its predecessor and 196% more than Astra. At $6 per million output tokens the per-task economics still work out, but anyone budgeting agent deployments by token volume should recalibrate. AA also measured the answer output speed at approximately 188 tokens/second for long prompts, with tasks averaging 7.1 minutes.
One more quality signal: Grok 4.7’s hallucination rate on AA-Omniscience fell to 29% from 34%, with accuracy broadly unchanged — so the added tokens are buying more self-checking, not more confabulation.
A new safeguard stack
Grok 4.7 ships with what SpaceXAI calls “an entirely new safeguard stack” — the strongest model it has tested on refusals and jailbreak resistance. In dual-use domains like cybersecurity and biology, the company says it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%. On HackerBench v0.3, SpaceXAI’s benchmark for risky and malicious cyber tasks, only 3.3% of risky dual-use prompts got through, while the model “rarely blocks legitimate security work.” The company has also begun giving select cybersecurity partners invite-only access to the model’s red-team capabilities for defense research.
Given how much skepticism rightly trails safety claims from any lab, the useful detail here is the shape of the numbers: high refusal on genuinely dangerous tasks, low refusal on legitimate ones. Most safety stacks trade one against the other; publishing both axes is at least a step toward accountability.
Availability
Grok 4.7 is available now in Cursor (all plans) and as the default model in Grok Build, plus the Grok API, OpenRouter, Vercel, and Cloudflare. A Grok 4.7 Fast variant doubles output speed at double the price, but runs only in Cursor and Grok Build — not the public API — and is excluded from Grok Build’s free tier. A US regional endpoint (https://us.api.x.ai/v1) keeps inference inside the United States at a 10% premium, and cache hits are discounted to $0.50 per million tokens.
Why this matters
Three takeaways. First, while OpenAI and Anthropic spent Tuesday fighting a pricing war measured in minutes, SpaceXAI demonstrated the other viable strategy: hold price, raise capability — and the benchmark table suggests it worked, at least for long-horizon professional work. Second, the specialisation pattern is getting sharper: Grok 4.7 doesn’t lead everywhere, but its leads (EEBench, Harvey legal) are in exactly the specialized professional domains where per-token prices compound fastest. Third, the AA token-usage numbers are the quiet headline: when a model more than doubles its token consumption to buy reliability, “price per million tokens” stops being the number that matters — price per completed task is, and that is a much harder thing to comparison-shop.