← All posts / Models

Tencent Open-Sources Hy4 Preview: A 770B MoE Flagship With 1M-Token Context

Tencent's Hy Team has released Hy4 preview, a 770B-parameter open-weight MoE model with 49B active parameters, a 1M-token context window, and benchmark wins over GLM-5.3 and Kimi K3.

Tencent Open-Sources Hy4 Preview: A 770B MoE Flagship With 1M-Token Context

Tencent’s Hy Team has open-sourced Hy4 preview, a new-generation flagship Mixture-of-Experts (MoE) large language model that lands squarely at the open-source frontier. Released on August 28, 2026 under the permissive Apache 2.0 license, the model ships with 770B total parameters, 49B activated per token, and a context window that stretches to a full 1 million tokens — a combination that until recently was the exclusive territory of closed, Western frontier labs.

The weights are available now on Hugging Face, ModelScope, GitCode, and CNB, in both a standard precision release and an FP8-quantized variant (Hy4 preview-FP8) aimed at practical multi-GPU deployments. For a “preview” release, the package is unusually complete: official vLLM and SGLang container images, a finetuning pipeline, and Tencent’s AngelSlim quantization toolkit are all published alongside the weights.

The Architecture: Sparse Attention, DeepSeek-Inspired

Beneath the headline numbers sits one of the more technically interesting open architectures of the year. Hy4 preview’s backbone consists of 78 layers, with the first layer using a standard dense feed-forward network and the remaining 77 layers using MoE. Each MoE layer contains 256 routed experts and 1 shared expert, and every token activates the top-8 routed experts plus the shared one — a sparsity pattern that keeps per-token compute at 49B parameters while the model as a whole carries 770B of capacity.

The attention module is where Tencent credits its lineage openly. Inspired by DeepSeek and GLM, Hy4 preview employs Gated DeepSeek Sparse Attention (Gated DSA) combined with IndexCache for cross-layer sparse index reuse. The residual pathway uses identity Hyper-Connections (iHC) with four residual streams to expand inter-layer information flow. In plain terms: rather than brute-forcing capability with dense compute, the model aggressively prunes what each token needs to look at — which is precisely how it can offer a 1M-token context window without the KV-cache memory bill of a dense model of comparable size.

There is also a native MTP (multi-token prediction) layer built into the model — 10B total parameters, 0.7B activated — designed for speculative decoding. In vLLM, enabling it is a single flag, and Tencent’s published serving recipe combines the MTP draft model with the FLASHMLA_SPARSE attention backend on an 8-way tensor-parallel setup. The FP8 weights make the whole thing servable on a single 8-GPU node class, which is exactly the hardware tier a well-funded startup or university lab can realistically rent.

The Benchmarks: Ahead of GLM-5.3 and Kimi K3

Tencent validated Hy4 preview with a blind side-by-side evaluation: 163 internal experts rated model outputs on 203 real engineering tasks. The results put Hy4 preview slightly ahead of the current Chinese open-weight leaders:

  • vs. GLM-5.3: 2.99 vs. 2.92 average score, with 46.8% wins / 12.8% ties / 40.4% losses
  • vs. Kimi K3: 2.99 vs. 2.94 average score, with 51.2% wins / 7.9% ties / 40.9% losses

Independent commentary cautions that the picture is nuanced — Kimi K3 reportedly still leads on GPQA Diamond, HLE, SWE-Marathon, and several pure reasoning benchmarks, making it the stronger mathematician and marathon coder. But on the aggregate blend of coding, reasoning, and agentic workloads that Tencent’s evaluation measures, Hy4 preview now holds the top average score among open-weight models.

The training philosophy behind those numbers is notable. Tencent says it partnered with practitioners inside the company — software engineers, game developers, finance analysts, and security experts — and built post-training data around the work those teams actually ship. The company claims four concrete productivity domains where the model pushes meaningfully further: long-horizon software engineering with better front-end visual taste; converting messy multi-file context into shareable documents, spreadsheets, and presentations; turning a single prompt into a playable game prototype; and scientific research spanning molecular dynamics, condensed matter physics, and pure mathematics.

Honest About Limitations

Refreshingly, the model card does not pretend the release is finished. Tencent lists known issues outright: the model spends longer than necessary reasoning through complex tasks, and it tends to over-verify its own work. The team frames the preview strategy explicitly — “we would rather ship early and hear what breaks” — pointing to the Hy3 preview cycle, where early community feedback directly shaped the substantially better final release.

That shipping philosophy has become a hallmark of the Chinese open-weight ecosystem: publish the weights, invite the criticism, iterate fast. It is also a pointed contrast with the deployment patterns of closed frontier labs, whose strongest models remain accessible only through metered APIs.

Why It Matters

Hy4 preview matters for three reasons. First, capability density: a 49B-active-parameter model that benches alongside GLM-5.3 and Kimi K3 means the open frontier continues to close the gap with closed models at an aggressive pace. Second, context economics: 1M tokens of context in an open-weight package — with sparse attention and FP8 weights to make serving tractable — unlocks long-document analysis, whole-repository code understanding, and long-horizon agent workflows for anyone with the hardware, not just anyone with an enterprise contract. Third, license freedom: Apache 2.0 with no territorial restrictions means commercial use, fine-tuning, and distillation are all on the table without negotiation.

The release also continues a striking 2026 pattern: the most capable openly downloadable models — Hy4 preview, GLM-5.3, Kimi K3, Qwen 3.8-Max — now all come from Chinese labs, released with permissive licenses while their American counterparts hold their frontier weights close. Whatever the strategic reasoning on each side, the practical effect is that the center of gravity for open-weight AI capability has shifted, and Hy4 preview is the current proof point.

For developers who want to try it, the barrier is lower than the parameter count suggests: pull vllm/vllm-openai:hy4-preview, load the FP8 weights with tensor-parallel 8, and the model serves an OpenAI-compatible API out of the box. Reasoning mode defaults to “high” for deep chain-of-thought on complex tasks and can be dialed to no_think for direct responses.

A full, non-preview Hy4 release is expected to follow as the team iterates on pre-training and post-training headroom. If the Hy3 trajectory is any guide, the final version will be worth waiting for — but there is little reason to wait to start building.