Nine Days Late, Now Live: Grok 4.7 Ships With Legal-Bench Bombshells and a $2 Price Tag
SpaceXAI's twice-delayed Grok 4.7 is finally here at $2/$6 per million tokens — nearly doubling Grok 4.6's Terminal-Bench score, crushing rivals on legal and electrical-engineering benchmarks, but still trailing Fable 5.1 on raw coding peak.
On September 2, Elon Musk promised Grok 4.7 in ten days. On September 11, he slipped the date. On September 21, it actually shipped. SpaceXAI’s flagship model is now live in Cursor and Grok Build, on the Grok API, and across third-party coding harnesses and model routers — and the launch card finally answers the question the entire delay drama raised: what was worth nine extra days of waiting?
The headline claim is value, not vanity. Grok 4.7 is “twice as fast, at half the price of comparable models,” served at the same price and speed as Grok 4.6: $2 per million input tokens and $6 per million output tokens. Against Fable 5.1 Max at $10/$50 or GPT-5.6 Sol Max at $4/$20, that pricing positions Grok 4.7 as the frontier’s discount tier. A fast variant doubles output speed at double the price.
A Bigger Base, Trained on Harder Problems
Under the hood, Grok 4.7 is a new, larger base model than Grok 4.6 — the widely reported 2.1 trillion parameters to Grok 4.6’s 1.5 trillion, following months of supplemental training on SpaceX engineering data. The launch notes describe a longer reinforcement learning run weighted toward “problems that take many hours to complete,” with specific gains in self-verification and long-context management. The model also now natively understands the Grok Bot harness, which SpaceXAI credits for improvements in conversational and general knowledge work.
That RL emphasis shows up most dramatically in agentic endurance. On Terminal-Bench 4.0, which measures multi-hour terminal work, Grok 4.7 xHigh scores 38.0% — nearly double Grok 4.6’s 20.3% and ahead of GPT-5.6 Sol Max at 37.3%, though well behind Fable 5.1 Max’s 57.9%. On DeepSWE v1.1, it posts 71.0% at high effort, within striking distance of GPT-5.6 Sol’s 72.7% and ahead of Grok 4.6’s 65.2%.
Where Grok 4.7 Actually Wins
The most startling numbers are outside pure coding. On the Harvey Legal Agent Benchmark, Grok 4.7 scores 19.6% — more than Grok 4.6 (15.8%), and several times GPT-5.6 Sol Max (2.5%) and Fable 5.1 Max (6.7%). On EEBench, an electrical-engineering evaluation, Grok 4.7’s 64.0% towers over Grok 4.6’s 53.0%, Fable 5.1’s 56.4%, and GPT-5.6 Sol’s 39.4% — a result that reads like the SpaceX engineering-data training paying off in domains with physical-systems reasoning.
On AA Briefcase v1.1, the multi-hour office-work benchmark, Grok 4.7 scores 1,657, ahead of Grok 4.6 (1,546) and GPT-5.6 Sol (1,487), trailing only Fable 5.1 (1,678). The GDPval Elo table tells a similar story: Fable 5.1 Max leads at 1735, Grok 4.7 xHigh sits close behind at 1695, Grok 4.6 at 1605, and GPT-6 Astra Max at 1542.
And where it loses, it loses honestly. CursorBench 4.0, which stresses longer-running coding tasks, puts Grok 4.7 at 46.3% — a clear step up from Grok 4.6’s 40.4% and ahead of GPT-5.6 Sol’s 41.7%, but behind Fable 5.1’s 51.8%. HealthBench Professional: 56.7%, improved over Grok 4.6’s 48.5% but behind GPT-5.6 Sol (60.5%) and Fable 5.1 (62.1%). Musk’s September promise that Grok 4.7 would “exceed all current models” did not materialize in absolute terms. What did materialize is arguably more commercially interesting: frontier-adjacent peak at commodity prices. SpaceXAI’s own launch chart concedes the framing, placing Grok 4.7 “at the frontier in price-performance” on CursorBench rather than at the frontier, full stop.
Independent trackers add a nuance the launch card doesn’t dwell on. Artificial Analysis scored Grok 4.7 xHigh at 46 on its Intelligence Index — but noted the model burns roughly 81k output tokens per task, more than double the 36k used by Grok 4.6 xHigh. Grok 4.7 thinks longer before it answers. At $6 per million output tokens that thinking is cheap; on latency-sensitive applications, it’s a real trade-off, and the new fast variant exists precisely to buy some of it back.
The Safeguard Stack Is the Sleeper Story
The quietest headline may matter most in enterprise procurement. Grok 4.7 ships with what SpaceXAI calls an entirely new safeguard stack — its strongest-tested model on refusals and jailbreak resistance. In dual-use domains it claims the lead on both benign utility and safe refusal: 62.4% on LatchBio’s biosafety benchmark, the top score, and on HackerBench v0.3, SpaceXAI’s benchmark for malicious cyber tasks, only 3.3% of risky dual-use prompts slip through while legitimate security work rarely gets blocked. The company has also opened invite-only red-team access for select cybersecurity partners doing defense research.
That’s a meaningful evolution for a lab once stereotyped as the loosenest of the frontier providers. If the numbers hold up under third-party scrutiny, safety calibration becomes a selling point rather than a compliance checkbox — and a direct answer to the industry-wide safety-standards conversations that OpenAI, Anthropic, and Google have reportedly been holding for weeks.
What It Means
Three takeaways for anyone picking models this week. First, the value tier got a genuine frontier occupant: for agentic coding and multi-hour task execution, $2/$6 with Grok 4.7’s scores is hard to beat on cost-per-completed-task. Second, specialized-domain strength is now a differentiator — the Harvey Legal and EEBench gaps aren’t rounding errors, they’re chasms, and they suggest training-data composition (rocket engineering records, in this case) creates defensible niches. Third, the peak of the peak still belongs to Fable 5.1, and if you need absolute maximum coding performance at 5–8x the price, the leaderboard still says pay for it.
The launch itself closes a messy chapter: a ten-day countdown, a public slip, and a ship date nine days late. But SpaceXAI has now released three model generations in roughly two months, each at the same aggressive price. Grok 4.7 doesn’t beat everything. It doesn’t have to. It just has to make everyone else’s pricing look uncomfortable — and on the evidence of this launch card, it does.