← All posts / Models

Ox Alpha: The Mystery Frontier Model Beating GPT-5.6 at Coding — and It's Free

An anonymous stealth model dubbed Ox Alpha appeared on OpenRouter this week with a 1M-token context window, 100 trillion free tokens per day, and coding scores that top GPT-5.6 — and fingerprinting evidence points straight at Zhipu's unreleased GLM-5.x.

Ox Alpha: The Mystery Frontier Model Beating GPT-5.6 at Coding — and It's Free

On August 20, 2026, an unbranded model called Ox Alpha quietly appeared on OpenRouter and OpenCode. No company name, no marketing page, no announcement thread. Just a model identifier (stealth/ox-alpha), a free tier, and a set of claims that read like a parody of frontier-LM launch posts: a 1,048,576-token context window, text/image/video input, and — according to the operator — capacity for 100 trillion tokens per day during a free promotional week.

Within 48 hours, the internet’s LLM community did what it always does with anonymous models: it started benchmarking, fingerprinting, and arguing about who built it. The early answers are striking. Ox Alpha isn’t just competitive with the frontier — on several coding and agentic evaluations, it appears to be ahead of it. And the growing pile of circumstantial evidence points toward one lab in particular: Zhipu AI, testing its unreleased GLM-5.x flagship in public under a pseudonym.

What Ox Alpha actually is

Strip away the mystery and the facts are simple. Ox Alpha is described by its operators as “a frontier model built for efficient coding, sustained agentic work, and real-world production use.” It accepts text, images, and video as input. Its context window is exactly 2^20 tokens — 1,048,576 — which is a suspiciously “engineered” number, the kind you pick when you want a clean power-of-two spec sheet rather than an organic one.

The free period is expected to run for roughly a week from launch, and the operators claim 99.99% uptime over the first days. ModelsAtlas, which tracks model listings across providers, shows Ox Alpha as released on August 20, 2026, with the same stealth/ox-alpha identifier and free access terms.

What makes the launch unusual is not the specs but the posture. Every major lab ships stealth models periodically — OpenAI, Google, and Anthropic have all used anonymous or codenamed deployments to gather real-world usage data before a public launch. But those tests are usually quiet, limited, and short-lived. Ox Alpha’s operators instead threw open the doors: 100 trillion tokens per day of capacity, open to anyone, with a stated agentic-coding focus. That’s not a cautious trial balloon. That’s a lab that wants a very large amount of production traffic, very fast.

The benchmark results

The numbers driving the speculation come from several independent sources.

On a 10-task deterministic agentic evaluation tracked by Orca Router, Ox Alpha reportedly passed 8 of 10 tasks (80% Pass@1), ahead of Claude Fable 5 (65%), GLM-5.3 and Grok 4.6 (62% each), and GPT-5.6-sol. On DeepSWE, a benchmark measuring sustained real-world software engineering performance, Ox Alpha scored 80% — beating both GPT-5.6 and Claude’s current flagships.

On Kingbench — a benchmark maintained by the creator community around OpenCode — the picture was slightly more modest: Ox Alpha scored 87.5%, landing second behind GLM-5.3 itself and ahead of Opus 4.8 and Qwen 3.8 Max.

That ordering is itself a clue. A stealth model that beats nearly everything except one specific commercial model — GLM-5.3 — looks less like a rival and more like a successor. If Ox Alpha is Zhipu’s next-generation GLM in disguise, the fact that it still trails (barely) the current GLM-5.3 on some community benchmarks while crushing it on others is exactly the pattern you’d expect from a pre-release candidate with incomplete post-training.

The fingerprinting case

The “who built it” question is where the story gets genuinely technical. Several independent analyses, most prominently by analyst Ben Davis, have assembled a fingerprinting case for Zhipu AI’s GLM lineage:

  • Tokenizer match. Multiple users report Ox Alpha uses the same tokenizer as the GLM model family — the single most diagnostic artifact in any LLM’s stack, since tokenizers encode years of training-data decisions.
  • Identical error messages. Reported (though not fully verified) shared error strings with GLM models — another deep-fingerprint artifact that’s hard to fake.
  • Stylistic and behavioral match. Reasoning traces, refusal patterns, and — per one widely-shared observation — the same audio-rejection behavior as GLM-5.3.
  • The launch pattern. Xiaomi’s MiMo team previously ran a similar stealth-then-reveal campaign, which is why some observers initially hedged between Zhipu and Xiaomi. But the tokenizer evidence tilts heavily toward Zhipu.

On Manifold Markets, a prediction market asking “Who is behind Ox Alpha?” has GLM-5.x as the dominant answer, with one trader summarizing the consensus bluntly: “99% sure it’s GLM-5.x, all the evidence points to it — same tokenizer, style matches, same audio rejection.”

It’s worth being precise about what this evidence can and cannot prove. Tokenizer reuse alone doesn’t confirm authorship — labs do license and reuse tokenizers. Stylistic matches in generated text can converge across models trained on similar corpora. And a stealth deployment could, in principle, be a deliberate decoy. But the convergence of four independent signals — tokenizer, error strings, behavioral quirks, and benchmark positioning relative to GLM-5.3 — makes the Zhipu hypothesis the clear frontrunner. No competing lab has produced comparable evidence.

Why labs do this

Stealth launches serve a specific purpose in the current AI market: they let a lab gather production-distribution data without committing to a name, a price, or a promise. If Ox Alpha is GLM-5.4 or GLM-5.5, Zhipu gets to observe how the model behaves under real agentic workloads — millions of genuine coding sessions, edge cases, failure modes — before putting its brand on the line. The 100-trillion-token free week is, in effect, the world’s largest unpaid beta test, with the community doing the QA for free.

There’s also a competitive-timing angle. GLM-5.3 shipped in mid-August to strong reviews — Interconnects’ analysis noted it “surpassed Moonshot AI’s Kimi K3” on many benchmarks, a significant marker for Chinese labs closing the frontier gap. If Zhipu already has a successor ready enough to test publicly, the cadence says something uncomfortable for Western labs: the iteration speed that once favored Silicon Valley may now be working against it.

The open questions

Three things remain unresolved:

  1. Who is actually behind it. The Zhipu hypothesis is strong but unconfirmed. The Manifold market and community consensus could still be wrong.
  2. What happens when the free week ends. A model with this profile — 1M context, video input, agentic coding focus — doesn’t stay free. Pricing will be the real signal of intent.
  3. Whether the benchmark lead holds. Stealth models often test above their eventual public scores, since pre-release candidates get cherry-picked checkpoints. GPT-5.6 and Claude Fable 5 weren’t standing still when Ox Alpha’s numbers were captured.

For now, Ox Alpha is the most interesting natural experiment in the LLM market: a frontier-class model, free to anyone, with no name attached. Whoever built it, the exercise demonstrates something the industry keeps relearning — that in 2026, anonymous frontier models draw more attention than branded ones, and that the community’s collective fingerprinting apparatus has become sharp enough that “stealth” is more marketing than concealment.

If Ox Alpha is Zhipu’s GLM-5.x, expect the reveal within weeks. If it isn’t, someone else just demonstrated frontier-class coding capability from complete anonymity — which might be the bigger story.