← All posts / Research

Two Steps, Not Ten: Independent 'Tauon' Optimizer Claims the Muon Crown on GPT-Mini

An independent researcher's Tauon optimizer — polynomial orthogonalization with spectral filtering — reportedly hits ~1.6 validation loss on GPT-Mini vs Muon's ~1.65 and AdamW's ~1.8, with ~8.5% faster steps. Another sign that optimizers are 2026's quiet battleground.

Two Steps, Not Ten: Independent 'Tauon' Optimizer Claims the Muon Crown on GPT-Mini

The most-upvoted AI story of this Sunday morning is not a model release, a funding round, or a regulation fight. It is an optimizer. An independent researcher posted “Tauon,” a new neural-network optimizer built on polynomial and orthogonalization ideas, and claims it beats Muon — the optimizer that frontier labs actually train on — on a small GPT-Mini transformer trained on TinyShakespeare, reaching roughly 1.6 validation loss versus Muon’s ~1.65 and AdamW’s ~1.8, with per-step time about 8.5% shorter than Muon. The post, surfaced at the top of AI Weekly’s live index for September 27, 2026, is spreading fast across machine-learning communities.

If those numbers hold up under replication, it matters — not because a TinyShakespeare run predicts frontier training, but because the optimizer has quietly become one of the highest-leverage decisions a lab can make.

What Tauon actually claims

According to the researcher’s description, Tauon attacks the optimization problem from a different direction than its predecessors. Where AdamW adapts per-coordinate learning rates from moment estimates, and Muon orthogonalizes the update matrix via Newton–Schulz iteration, Tauon leans on polynomial and orthogonalization techniques with two named mechanisms:

  • Spectral filtering — selectively damping components of the update by their spectral profile, so that the optimizer’s effort concentrates on the subspaces where gradient information is actually reliable.
  • Coefficient scheduling — treating the polynomial coefficients of the update rule as scheduled quantities rather than fixed hyperparameters, which the author says reduces the effective optimization to two steps per iteration.

The headline results on the small GPT-Mini benchmark: Tauon ~1.6 validation loss, Muon ~1.65, AdamW ~1.8 — with a per-step wall-clock cost roughly 8.5% lower than Muon’s. On a speedrun benchmark where Muon itself already dominates AdamW, that is a double win: fewer steps to a better loss, and cheaper steps.

The obvious caveats deserve stating plainly. This is a single independent researcher’s post, not a peer-reviewed paper with ablations. GPT-Mini on TinyShakespeare is a toy by frontier standards — millions of parameters, character-level Shakespeare, minutes of compute. The graveyard of optimizers that won at small scale and diverged at large scale is large. And 2026 has already produced a pointed reminder that Muon’s success may not come from the precise geometric story its fans tell: a May 2026 analysis (“Muon is Not That Special: Random or Inverted Spectra”) argued that the optimizer family’s gains survive even when the spectral preconditioning is deliberately distorted, suggesting some of the benefit is more mundane than the theory implies.

Why an optimizer post trends: Muon ate the frontier

To understand why a Reddit optimizer post can outrank billion-dollar deal news for a news cycle, look at what Muon did over the past two years.

Muon, published by Jeremy Bernstein in late 2024, was initially a speedrun curiosity — the optimizer behind NanoGPT and CIFAR-10 training records. The serious turn came in February 2025, when Moonshot AI showed Muon was scalable to large-language-model pretraining, then put its weight behind the claim by training and open-sourcing Moonlight, a 3B/16B MoE model, entirely with an improved Muon variant.

Then came the proof at scale. Kimi K2 — a trillion-total-parameter, 32B-activated MoE — was pretrained on 15.5 trillion tokens using MuonClip, a Muon extension pairing the optimizer with a QK-clip technique to kill attention-logit instabilities. Moonshot reported the entire run went through without a single loss spike, an almost boastful claim for anyone who has babysat a trillion-token job. Reporting around Moonshot’s deployments, including a PyTorch blog on running Muon with DeepSpeed, has circulated training speedups over AdamW in the 25–35% range — figures that, at frontier scale, translate into hundreds of millions of dollars of compute per training run.

The trend only deepened in 2026. Kimi K3’s training stack reportedly applies Muon per attention head for stability. Alibaba’s Qwen team has described training recipes that assign Muon and AdamW to different weight categories, treating optimizer choice as an architecture-level decision rather than a hyperparameter afterthought. When the world’s most compute-hungry labs are all converging on one optimizer from 2024’s speedrunning scene, the frontier slot for “what beats Muon” is genuinely open — and valuable.

The stakes: efficiency is the only free lunch left

Training economics in 2026 are brutal. Akamai just signed an $11.6 billion, seven-year cloud commitment with Anthropic; Brookings projects the US AI infrastructure buildout at $10.3 trillion through 2032 — the largest single-industry buildout as a share of GDP in American history. In that world, an optimizer that reaches the same loss in fewer, cheaper steps is not an academic nicety. It is leverage on the single biggest line item in every frontier lab’s budget.

This is why the optimizer niche keeps producing viral results: Keller Jordan’s Modded-NanoGPT speedruns (where Newton-Muon variants now trade step-count records with NorMuon), the Kaon critique papers, and now Tauon. Each one is a bid for the same prize — a measurable, replicable efficiency edge on the path everyone must walk.

Tauon’s “two optimization steps” framing, if it survives scrutiny, would fit neatly into that economy. Newton–Schulz orthogonalization — Muon’s core loop — costs several matmul-heavy substeps per iteration. An approach that achieves comparable conditioning with two steps, at 8.5% lower per-step cost, attacks exactly the overhead that Muon added in exchange for its stability.

What would make it real

For Tauon to graduate from viral post to training stack, the community will want to see, roughly in order:

  1. Independent speedrun replications on Modded-NanoGPT-style benchmarks, where Muon’s records are continuously verified.
  2. Scaling behavior — the Kimi-style test: does the advantage survive a multi-billion-parameter MoE run, or does the polynomial machinery diverge where AdamW’s dull robustness glides through?
  3. A paper with ablations separating the contribution of spectral filtering from coefficient scheduling, ideally against the Kaon critique’s finding that distorted spectra barely hurt.
  4. Interaction with stability tricks — whether it composes with QK-clip-style fixes the way MuonClip did.

None of that has happened yet. What has happened is that on a slow Sunday in late September, the single most-engaged story in AI tracking circles is a grassroots optimizer result — a healthy reminder that for all the trillion-dollar infrastructure headlines, the field’s most consequential open question is still the oldest one: what is the best way to descend a loss surface?

The sources for this story are rendered below.