← All posts / Models

Tencent Open-Sources Hy4 Preview: A 770B-Parameter MoE Frontier Model with 1M-Token Context

Tencent's Hunyuan team releases Hy4 preview under Apache 2.0: 770B total parameters, 49B active per token, a 1M-token context window, and benchmark scores that edge out GLM-5.3 and Kimi K3 on real-world engineering tasks.

Tencent Open-Sources Hy4 Preview: A 770B-Parameter MoE Frontier Model with 1M-Token Context

On August 28, 2026, Tencent released and open-sourced Hy4 preview, the newest flagship large language model from its Hunyuan team — and it is a serious statement of intent. With 770 billion total parameters, 49 billion activated per token, and a context window exceeding one million tokens, Hy4 preview lands squarely at the open-source frontier, directly challenging DeepSeek, Zhipu’s GLM-5.3, Moonshot’s Kimi K3, and Alibaba’s Qwen family for leadership of the open-weight race.

The weights are available now under the fully permissive Apache 2.0 license, in both BF16 and FP8 variants, on Hugging Face, ModelScope, GitCode, and CNB. For teams that prefer not to self-host, the model is also served through Tencent’s WorkBuddy and CodeBuddy products (free for two weeks at launch), Yuanbao, ima, and via API through Tencent Cloud TokenHub and OpenRouter.

What’s inside the architecture

Hy4 preview is a next-generation Mixture-of-Experts (MoE) model, and the spec sheet reads like a survey of the most effective ideas in recent open research:

  • 78 layers, where the first layer uses a standard dense FFN and the remaining 77 layers use MoE — each containing 256 routed experts plus 1 shared expert, with every token activating the top-8 routed experts alongside the shared one.
  • A native MTP (multi-token prediction) layer — 10B total parameters, 0.7B activated — built in for speculative decoding, which dramatically accelerates inference.
  • Attention based on Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache for cross-layer sparse index reuse. Tencent explicitly credits inspiration from DeepSeek and GLM here — a telling example of how quickly architectural innovations now propagate across the Chinese open-source ecosystem.
  • Residual pathways using identity Hyper-Connections (iHC) with 4 residual streams to expand inter-layer information flow.
  • A 6144 hidden size, 64 attention heads, KV compression down to 512 dimensions, and a 120,832-entry vocabulary.

The headline capability for practitioners is the 1M-token context window. Combined with the sparse-attention stack, that positions Hy4 preview for exactly the kind of long-horizon work — whole-codebase reasoning, multi-document analysis, month-long agent sessions — where frontier labs are increasingly competing.

Built for productivity, not just benchmarks

Tencent frames Hy4 preview as a model “built for productivity.” Rather than optimizing purely for leaderboard wins, the Hunyuan team partnered with internal experts — software engineers, game developers, finance analysts, and security specialists — and built training data around the work those teams actually ship. The claimed gains cluster in four areas:

  1. Software engineering: stronger understanding, planning, debugging, and verification of long-horizon development tasks, plus improved visual taste in front-end work.
  2. Office and analysis: turning messy context spread across many files into shareable artifacts — documents, spreadsheets, presentations — with more precise data analysis, equations, and financial models.
  3. Game development: generating a playable prototype from a single prompt and working fluently with game engines across multi-turn refinement.
  4. Scientific research: progress on AI research, molecular dynamics, condensed matter physics, and pure mathematics.

The numbers

The benchmark story is competitive at the top. On GPQA Diamond, Hy4 preview scores 92.3. On SWE-bench Pro (public) it reaches 65.7, and on SWE-bench Multilingual it hits 82.9 — ahead of GLM 5.3 (81.3), Kimi K3 (80.8), and DeepSeek V4 Pro (77.3) on that last test. On the Humanity’s Last Exam–style HLE set it posts 55.40.

Tencent also ran a blind side-by-side internal evaluation: 163 experts graded 203 engineering tasks, and Hy4 preview came out slightly ahead of both GLM-5.3 (2.99 vs. 2.92 average out of 4.00; 46.8% wins / 12.8% ties / 40.4% losses) and Kimi K3 (2.99 vs. 2.94; 51.2% wins / 7.9% ties / 40.9% losses). Community analyses caution the picture is nuanced — against Kimi K3, Hy4 trails on 7 of 12 external benchmarks including Terminal-Bench — but the overall claim that Hy4 preview sits in the top tier of open-source models holds up.

The standard caveat applies: the blind evaluation was conducted and graded internally by Tencent. Independent verification will come from the community, and the Apache 2.0 release makes that scrutiny possible in a way closed APIs never allow.

Shipping early, on purpose

The Hunyuan team is candid about limitations: Hy4 preview sometimes spends longer than necessary reasoning through complex tasks and tends to over-verify its own work. Their stated philosophy is to “ship early and hear what breaks” — the same approach they credit for making Hy3 substantially better over time.

Deployment is well-supported from day one. Official prebuilt containers exist for both vLLM (vllm/vllm-openai:hy4-preview, with an 8-way tensor-parallel FP8 recipe and MTP speculative decoding enabled) and SGLang (lmsysorg/sglang:hy4-preview, multi-arch for x86 and Arm). Reasoning effort defaults to a deep chain-of-thought “high” mode, with an explicit no_think escape hatch for direct responses. Tencent’s AngelSlim toolkit handles further compression and quantization.

Why it matters

Three bigger threads make this release more than a routine model drop.

First, the open-weight frontier keeps moving to China. Between DeepSeek’s pending ~$7.4B round at a ~$74B valuation, GLM-5.3, Kimi K3, Qwen 3.8, and now Hy4 preview, the densest concentration of frontier-class open models is Chinese — at the same moment Washington is drafting rules to close the “remote GPU rental” loophole and Stanford’s AI Index reports the US–China capability gap has effectively closed.

Second, architecture is consolidating in the open. Hy4’s Gated DSA attention and MTP speculative decoding are direct descendants of techniques popularized by DeepSeek. Open weights mean good ideas diffuse in months, not years — and every lab builds on the last one’s wins.

Third, self-hosting a frontier model is now practical. An FP8-quantized 770B MoE that activates only 49B parameters per token, served with speculative decoding on an 8-GPU node, is within reach of a well-funded enterprise or research lab — no API dependency, no rate limits, full data control, under a license with no restrictions.

Hy4 preview is a preview, with real headroom left in both pre-training and post-training. But as a marker of where the open-source frontier stands at the end of August 2026 — massive, sparse, long-context, productivity-focused, and fully permissive — it is hard to ignore.