Bigger, Cheaper, Stubborner: Inside xAI's Grok 4.7
xAI's Grok 4.7 pairs a larger base model with unchanged $2/$6 pricing, nearly doubling Terminal-Bench scores and topping electrical-engineering and biosafety benchmarks.
xAI shipped Grok 4.7 on September 21, 2026, and the release reads like a deliberate statement about where the frontier-model market now stands. The company describes it as its “most capable model for coding and knowledge work,” and the headline pitch is almost paradoxical: a new, larger base model — reportedly around 40 percent more weights than Grok 4.6 — served at exactly the same price and speed as its predecessor. In a year where Anthropic’s Fable 5.1 charges $10 per million input tokens and $50 per million output tokens, Grok 4.7 holds the line at $2 and $6. Twice as fast, at half the price of comparable models, is the claim xAI puts right under the title.
What actually changed
Under the hood, Grok 4.7 is not a fine-tune or a repackaged checkpoint. According to the official announcement, xAI trained it with a longer reinforcement learning run on a harder mix of tasks, deliberately weighted toward problems that take many hours to complete — the multi-hour coding sessions, terminal marathons, and document-production grind that agentic workflows increasingly demand. The company also says the model is better at verifying its own work and managing longer context, and that it trained Grok 4.7 to natively understand the Grok Bot harness, which sharpens conversational tasks and general knowledge work.
The context window remains 500K tokens as on Grok 4.6, and pricing has an interesting wrinkle beyond the base rates: cached input runs $0.50 per million tokens, and xAI also serves a fast variant with twice the output speed at twice the price. That two-tier structure is a quiet acknowledgment that in 2026, latency has become a first-class product feature alongside raw intelligence.
The benchmark picture
The numbers xAI published tell a story of uneven but genuine progress. On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 scores 46.3% — a solid jump from Grok 4.6’s 40.4%, and clearly ahead of GPT-5.6 Sol’s 41.7%, though still behind Anthropic’s Fable 5.1 at 51.8%. DeepSWE v1.1 lands at 71.0% (marked as a high-effort score), essentially at parity with GPT-5.6 Sol’s 72.7% and Fable 5.1’s 70.0%.
The most dramatic gains come in exactly the domains the training mix was weighted toward. Terminal-Bench 4.0, which measures multi-hour terminal work, jumps from 20.3% on Grok 4.6 to 37.6% — nearly doubling, and edging past GPT-5.6 Sol’s 37.3% (Fable 5.1 still leads at 57.9%). EEBench, an electrical-engineering evaluation, leaps from 53.0% to 64.0%, far ahead of GPT-5.6 Sol’s 39.4% and Fable 5.1’s 56.4%. On the Harvey Legal Agent Benchmark, Grok 4.7 improves from 15.8% to 19.6%.
Independent measurement broadly corroborates the picture. Artificial Analysis scored Grok 4.7 two points higher than Grok 4.6 on its Intelligence Index, with particularly strong performance on agentic knowledge-work tasks — while noting the gains concentrate in long-horizon, tool-using work rather than short-form chat.
Professional knowledge work shows the same pattern. On GDPval, an Elo-rated evaluation where AI handles tasks done by lawyers, nurses, and financial analysts, Grok 4.7 scores 1695 at xhigh effort, up from Grok 4.6’s 1605 and ahead of GPT-6 Astra’s 1542, though Fable 5.1 still tops the table at 1735. AA Briefcase v1.1, another multi-hour office-work benchmark, shows a similar step up from 1,546 to 1,657.
A new safeguard stack
Perhaps the most consequential part of the release is buried in the safety section. Grok 4.7 ships with what xAI calls an entirely new safeguard stack, and the company claims it is the strongest model they have tested on refusals and jailbreak resistance. The specifics are notable: Grok 4.7 tops LatchBio’s biosafety benchmark at 62.4%, and on HackerBench v0.3 — xAI’s benchmark for risky and malicious cyber tasks — it allows only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. That balance of “high capability, low over-refusal” is the needle every lab is trying to thread in dual-use domains like cybersecurity and biological research. xAI has also begun giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.
This matters beyond xAI. Frontier-lab competition has shifted from pure benchmark bragging rights toward who can ship the most capable model that still refuses the dangerous stuff — regulators, enterprise buyers, and the labs’ own safety teams all increasingly score both sides of that ledger.
Availability
Grok 4.7 is available now in Cursor and Grok Build, through the Grok API, third-party coding harnesses, and model routers and cloud platforms. Getting started is as simple as curl -fsSL https://x.ai/cli/install.sh | bash or a free spin in Grok Build.
Analysis: the price-performance squeeze
Strip away the benchmark tables and the strategic move is clear. xAI is contesting the frontier not by claiming the single highest score on any one eval — Fable 5.1 still holds the coding crown and GPT-6 Sol wins some rows — but by holding near-frontier performance at roughly one-fifth the output price of Anthropic’s flagship. On xAI’s own price-performance chart for CursorBench 4.0, Grok 4.7 sits at the frontier: 46.3% at $6-per-million output is a different value proposition than 51.8% at $50.
That is a squeeze on everyone above xAI’s price point. When a 40-percent-larger model ships at unchanged prices and nearly doubles on agentic terminal work, the effective cost of “good enough frontier” keeps falling — and the labs charging premium rates have to justify the delta with either capability gaps (Fable 5.1’s coding lead) or ecosystem depth. With OpenAI shipping GPT-6 Sol and Luna the day before Grok 4.7’s release, and Anthropic’s Claude Opus 5.5 landing the same week, September 2026 is shaping up as the most crowded frontier-model window yet. xAI’s bet is that in that crowd, the sticker price is what people remember.