← All posts / Models

Named Before It Ships: Musk Reveals Grok 4.8, a 2.5T Model on a Rewritten C++ Stack, While 4.7 Is Still Missing

In a single reply post, Elon Musk named Grok 4.8 — 2.5 trillion parameters, trained on SpaceXAI's rewritten C++ stack — while Grok 4.7 remains unreleased past its fourth missed date. We trace the four-month paper trail behind the claim.

Named Before It Ships: Musk Reveals Grok 4.8, a 2.5T Model on a Rewritten C++ Stack, While 4.7 Is Still Missing

At 9:24 PM Eastern on September 13, someone on X asked Elon Musk what SpaceXAI had planned for the month beyond Grok 4.7. The reply named a model that, until that moment, did not exist in public: “Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RL.”

That single sentence is now the entire public record for Grok 4.8. No model page, no API identifier, no price, no benchmark, no release note — as of September 14, xAI’s newsroom, model directory, and developer documentation contain no mention of either 4.8 or 4.7. The company’s newest released model remains Grok 4.6, which shipped August 12. What makes the post notable is not the number but the context: Grok 4.7 is itself still unreleased, having missed its September 11 target, with Musk saying only that it needs “a few more days.” SpaceXAI has now named the successor to a model it has not yet shipped.

The post, and what it actually claims

Musk’s reply packs three substantive claims into one line. First, scale: 2.5 trillion parameters, which would make Grok 4.8 roughly 67% larger than the 1.5T Grok 4.5/4.6 pair and about 19% larger than the 2.1T Grok 4.7 that is still waiting in the queue. Second, infrastructure: the model is said to be trained on SpaceXAI’s new C++ software stack — the rewritten training and inference pipeline the company has been building since May. Third, timing: pre-training finishes “this week,” with reinforcement learning to follow.

None of this comes with a spec sheet. xAI has never published a parameter count for any Grok 4-series model; every figure in the ladder — 0.5T for the production model in May, 1.5T for V9-Medium (which became Grok 4.5), 2T then 2.1T for what became 4.7, and now 2.5T — traces back to Musk’s own posts. And the founder’s mapping of counts to names has already moved once: a July 18 post attached a 2T model to “Grok 4.6?” with a question mark, only for the 4.6 label to land ten days later on the 1.5T model, with the 2T-class run renamed 2.1T and reassigned to 4.7.

There is also a technical caveat that parameter headlines tend to bury. Grok models are mixture-of-experts systems, and total parameter count describes memory footprint more than throughput. Musk himself noted that Grok 4.7 would be “slightly slower to serve” than 4.6 despite being larger, for exactly this reason. A 2.5T model on a faster stack could serve faster or slower than a 2.1T model on the old one — nobody outside xAI can currently say which.

Four months of C++ promises, one sentence of payoff

The most consequential part of the September 13 post may be the “C++ software stack” mention, because it is the first time the rewrite has been attached to a numbered model. The paper trail runs four months, all from one account:

  • May 28 — SpaceX has “almost finished” V1.0 of an in-house AI training stack in C that “exact-maps to 220k GB300s with 800G NICs,” heavy on pipeline parallelism, “getting as close to bare metal as possible.”
  • May 28 (later) — Next: an inference stack in C “for simultaneous high-speed RL across a large block of GB300s,” with the admission “(We do use a little C++ tbh, but not much).”
  • June 29 — “Truly massive gains will come in ~3 months when the entire training and inference stack is written in C/C++ and massively simplified (most software layers will be deleted completely).”
  • July 8 — Grok 4.5 is “not yet” on the new inference software; “doubling or more of the current speed is probably achievable.”

Three months from June 29 lands in late September, so a model finishing training on the new stack in mid-September fits inside Musk’s own stated window — arguably the first time a deadline-shaped claim from this timeline has aligned with reality. The strategy itself is a bet that software overhead, not silicon, is the binding constraint at Colossus-scale: delete the layers, map the code directly onto GB300 hardware, and recover speed that generic frameworks leave on the table. If the stack works as described, it functions as a capacity multiplier across every future Grok run, which is why one line about a software stack may matter more than the parameter count above it.

What “finish training this week” has meant before

The only honest calibration for the new claim is Musk’s own track record with sentences of this shape, and it argues for patience:

  • Grok 4.5: “finished training” posted May 25, released July 16 — 52 days later. (The same post promised “2 to 3 weeks to public release.”)
  • Grok 4.6: “will finish training next week” posted July 18, released August 12 — 25 days later.
  • Grok 4.7: “initial training is complete” posted August 12 — still unreleased as of September 14, 33 days and counting, with the September 11 launch target missed and an RL setting blamed for the delay.

Apply those gaps to a mid-September training finish and Grok 4.8 lands somewhere between early October and early November — if it follows either precedent at all. Every Grok 4 release so far has slipped its first stated date; the shorter end of Musk’s own estimates has not once held.

What this means for the queue

The awkward question is what happens to Grok 4.7. SpaceXAI has recent precedent for short-lived flagships — Grok 4.5 shipped July 16 and was superseded 27 days later — but 4.7 now risks a stranger fate: shipping as a stopgap weeks before its own successor, or being skipped entirely, with the wait folded into 4.8. Nothing in the September 13 post answers that. The reply that named 4.8 did not mention 4.7 once.

For anyone tracking releases rather than announcements, the confirmation signals to watch are concrete: a 4.8 model page or release note on x.ai or docs.x.ai; a post saying RL has started (which dates the clock); any movement on 4.7, since every 4.8 date depends on whether its predecessor ships first; and pricing relative to Grok 4.6’s $2/$6 per million input/output tokens — a larger model on a faster stack could move that number in either direction.

The pattern behind the number

Strip away the specifics and the Grok 4.8 reveal continues a recognizable playbook: announce through reply posts, attach big round numbers, and let the timeline stay fuzzy. The 2.5T figure extends a ladder the founder has been climbing publicly since May, and the C++ stack story turns an infrastructure project into a product feature. None of it is false, and the engineering bet — bare-metal software mapped to a single hardware platform — is genuinely interesting. But as of today, Grok 4.8 exists in exactly one place: a reply on X, posted while the model before it remains unshipped. The distance between that sentence and a launchable product is measured in Musk’s own historical gaps: 25 days, 52 days, or still counting.