← All posts / Models

Ox Alpha: The Anonymous Stealth Model Nobody Will Claim — and Everyone Is Using

An unidentified frontier model called Ox Alpha appeared free on OpenRouter with a 1M-token context window, beat GPT-5.6 and Claude on community coding benchmarks, and tokenizer fingerprinting now points to one surprising suspect.

Ox Alpha: The Anonymous Stealth Model Nobody Will Claim — and Everyone Is Using

On August 20, 2026, a new model quietly appeared on OpenRouter under the identifier stealth/ox-alpha. No company. No model card. No marketing blog post. Just a listing that described “a reasoning model designed for coding, sustained agentic work, and production workloads” — and a price tag of zero. Within seventy-two hours, it was one of the most talked-about models in the AI community, the subject of a Bloomberg story, a live prediction market on its creator’s identity, and a genuine benchmark mystery.

What Ox Alpha actually is

Strip away the mystique and the listing’s technical claims are unusually concrete:

  • A 1,048,576-token context window — a full 1M tokens, putting it in the same tier as the largest context offerings from Google and Meta.
  • Multimodal input across text, images, and video — a combination still rare outside the biggest frontier labs.
  • Positioning as a long-horizon agentic model: the description emphasizes sustained coding work and production workloads, not chat.
  • Free access — OpenCode, which distributes the model alongside OpenRouter, has said the free window lasts roughly a week.

There are still no official benchmarks. The provider has published nothing. Everything the community knows about Ox Alpha’s capability comes from crowd-sourced testing during the free window — and that testing has been surprisingly flattering.

It beats the big names on community benchmarks

Independent users running community benchmarks during the free window have reported results that would be headline-worthy from any named lab. On DeepSWE, a long-horizon software engineering benchmark, community runs put Ox Alpha at roughly 80% — ahead of GPT-5.6 Sol and Claude Fable 5, two of the strongest coding models money can buy. One widely shared analysis reported the model maintaining a clean pass across 51,469 regression tests during sustained code-generation work — the kind of consistency that matters for real refactoring jobs, not just benchmark sprint tasks.

Real-world usage has been more mixed, which is expected for a free, presumably capacity-limited preview. Users have reported throughput dropping to 20–30 tokens per second during peak congestion, occasional dropped connections, and inconsistent agentic behavior under load. In other words: the raw capability appears frontier-grade, but the serving infrastructure is clearly not provisioned like a commercial launch. That pattern — top-tier model, ramshackle delivery — is itself a clue.

The fingerprint: +75 tokens, every time

The identity hunt is where Ox Alpha becomes genuinely interesting. Since no one will claim the model, the community turned to black-box forensics.

The sharpest result came from tokenizer fingerprinting. Because every lab builds its own tokenizer, the way a model splits pinned strings of CJK text, emoji, and edge-case Unicode acts as a signature. One developer ran Ox Alpha against 25 test prompts and compared token counts against known models. The result: Ox Alpha’s token counts matched Z.ai’s GLM-5.3 exactly — with a constant offset of +75 tokens on every single text. A constant offset across every prompt almost certainly means the same tokenizer behind a hidden 75-token system prompt.

Earlier speculation had cycled through other candidates, and a Manifold Markets prediction market on “Who is behind Ox Alpha?” is still trading. But the tokenizer evidence, combined with Techstrong’s reporting that the model “could be from China,” has concentrated the consensus on Zhipu AI (Z.ai) testing an unreleased GLM-5.x flagship in the wild — possibly a variant tuned for coding and agentic work, possibly something new entirely. Z.ai has not commented, which of course proves nothing and fuels everything.

If true, it would be a striking inversion of the old pattern. For years the “mystery model” genre — from gpt2-chatbot onward — has been the signature move of American labs road-testing frontier models before launch. A Chinese lab running the same playbook on a US-centric router, at a moment when open-weight Chinese models are already dominating global usage rankings, would say a great deal about how confident these labs have become.

The catch: your prompts are the product

Here is the part every developer pasting code into Ox Alpha this week should read carefully. The OpenRouter Stealth Program agreement states that user content may be collected, retained, and provided to the anonymous providers — in plain terms, your prompts and completions can be kept and used for training by a company you cannot name. TechTimes framed it bluntly: the model retains every developer prompt, and you have no idea who is holding them.

There is one meaningful nuance: on OpenCode specifically, Ox Alpha is served under zero data retention — stricter than the lowest-retention tiers OpenAI, Google, or Anthropic offer on their own platforms. So the privacy posture depends entirely on which surface you use. On OpenRouter’s stealth listing, assume everything you send can be retained by an unidentified party. On OpenCode, it is not.

The practical guidance writes itself: enjoy the free frontier compute, but do not paste passwords, API keys, client files, or proprietary code during the free window — at least not on the retaining surface.

Why labs do this

Stealth releases are not charity; they are a strategy with three well-understood payoffs. First, clean benchmark data: a free, anonymous model attracts heavy organic traffic, giving the provider thousands of real-world agentic workloads to evaluate against — far better signal than internal evals. Second, hype without liability: if the model stumbles, no brand is damaged, because there is no brand. If it shines, the eventual reveal lands as a triumph. Third, competitive reconnaissance: watching how your model ranks on community leaderboards against named rivals, before committing to a launch date, is information money cannot easily buy.

The roughly one-week free window fits that playbook precisely. Whenever the window closes, expect one of two things: the model disappears and its learnings surface in a future named release, or the operator steps forward and converts the buzz into a commercial launch.

Bottom line

Ox Alpha is the most interesting non-announcement of the month: frontier-class coding and agentic performance, a 1M-token multimodal context, a price of zero, and a creator that fingerprinting evidence ties to Z.ai’s GLM lineage. Whether it turns out to be a stealth benchmark run for an unreleased flagship or a deliberate demonstration that the frontier can be served anonymously, it marks a shift worth watching — the mystery-model genre is no longer an American lab exclusive. Use the free window; just remember that with an anonymous provider, the free tier is you.

Try it at your own risk at openrouter.ai/stealth/ox-alpha — and mind what you paste.