← All posts / Models

Tencent Open-Sources Hy4 Preview: 770B MoE Flagship With 1M Context and a 31.8% Self-Tuning Speedup

Tencent Hunyuan's new open-weight flagship packs 770B total / 49B active parameters, a 1M-token window, and an early recursive self-improvement loop that lifted its own inference throughput 31.8%.

Tencent Open-Sources Hy4 Preview: 770B MoE Flagship With 1M Context and a 31.8% Self-Tuning Speedup

Tencent’s Hunyuan team released and open-sourced Hy4 preview on August 28, 2026, a next-generation flagship large language model that pushes the lab’s open-weight lineup into a new size class: 770 billion total parameters with only 49 billion activated per token, and a context window that exceeds 1 million tokens. The weights are live on Hugging Face and ModelScope in both BF16 and FP8 variants, and the model is available globally through Tencent’s WorkBuddy and CodeBuddy products, Yuanbao, ima, and via API through Tencent Cloud TokenHub and OpenRouter.

The release matters for three reasons. First, it escalates the open-weight arms race — Hy4 preview enters a tier previously occupied by the largest open models from DeepSeek, Zhipu, and Moonshot, and Tencent claims it lands “among the top tier of open-source models.” Second, its headline evaluation isn’t a public benchmark but an internal blind human study, an unusual and arguably more honest signal. Third, and most striking: Tencent says the model participated in its own development, establishing an early-stage recursive self-improvement loop that measurably sped up its own serving stack.

Architecture: DeepSeek-inspired sparse attention at 770B scale

Hy4 preview is a Mixture-of-Experts transformer with 78 layers. The first layer uses a dense feed-forward network; the remaining 77 layers each carry 256 routed experts with top-8 routing plus one shared expert. The per-token activation of 49B out of 770B total gives it a roughly 6.4% activation ratio — aggressive sparsity that keeps inference costs in check despite frontier-class total capacity.

The attention module is where the DeepSeek lineage shows. Hy4 preview employs Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache, a mechanism for cross-layer sparse index reuse that reduces the memory and compute burden of long-context attention. The residual pathway uses identity Hyper-Connections (iHC) to widen information flow between layers. A native multi-token prediction (MTP) layer adds 10B total parameters (0.7B activated) for speculative decoding — vLLM ships a recipe that serves the FP8 model with MTP on 16×B200 or 8×B300 GPU nodes using FLASHMLA_SPARSE attention and the new hy_v4 tool-call and reasoning parsers.

That 1M-token context (1,048,576 tokens precisely) is a fourfold jump over the 256K window of Hy3 preview, Tencent’s 295B/21B model from July, and puts Hy4 preview in the same context league as the longest-window frontier models.

Trained for productivity, judged by experts

Tencent positions Hy4 preview as a productivity model rather than a chatbot, with emphasis on coding, office work, data analysis, game development, and scientific research. The training data was co-created with Tencent domain experts in software engineering, gaming, finance, and security, and the model was co-designed with the teams behind WorkBuddy and CodeBuddy.

The headline evaluation reflects that focus. In an internal blind evaluation with 163 experts grading 203 engineering tasks, Hy4 preview scored an average of 2.99 out of 4.00 — slightly ahead of GLM-5.3 (2.92) and Kimi K3 (2.94), the two Chinese open-weight flagships it directly competes with. Human expert panels are slower and smaller than automated leaderboards, but they resist the contamination and gaming that plague public benchmarks — a point of growing sensitivity across the industry.

On capability specifics, Tencent highlights stronger understanding, planning, debugging, and validation for long-context development tasks, improved front-end generation quality, deeper financial analysis, and full workflow support from information processing through document, spreadsheet, and presentation creation. In game development, the model can reportedly generate a playable prototype from a single natural-language request and then refine complex projects over multi-turn interactions. In research, the company cites gains in AI R&D, molecular dynamics simulation, condensed-matter physics, and fundamental mathematics.

The recursive self-improvement loop

The most technically interesting claim in the announcement concerns Hy4 preview’s role in its own build. For the first time, the model participated in automated optimization of its training methods, data strategies, evaluation frameworks, and low-level operators — proposing approaches, running experiments, and iterating on the results, with the generated code, logs, and feedback feeding into subsequent rounds of exploration. Tencent describes this as an early-stage recursive self-improvement loop.

The loop extends to inference. Hy4 preview autonomously analyzed bottlenecks in its serving system and ran multiple optimization rounds on operator fusion and communication, lifting end-to-end throughput by 31.8% over baseline, with gains holding across different context lengths and concurrency levels. A model that tunes the infrastructure it runs on — and does so effectively — is a meaningful data point for the “AI accelerating AI development” thesis that every frontier lab is now chasing.

Pricing and availability

Hy4 preview is free on WorkBuddy and CodeBuddy for a two-week launch window, and free access to the older Hy3 on both platforms has been extended through September 30. API pricing undercuts most frontier closed models by a wide margin: $0.834 per million input tokens, $2.501 per million output tokens, and $0.042 per million cached tokens through Tencent Cloud TokenHub and OpenRouter.

The “preview-first” release strategy is deliberate: Hunyuan ships early, gathers real-world feedback through official releases, and folds lessons into the next iteration. Tencent says the next batch of Hy4-series models is expected soon.

What it means

Hy4 preview lands in a week when China’s AI token usage topped 500 trillion per day and open-weight releases from Alibaba, Zhipu, Moonshot, and now Tencent are arriving at a cadence Western closed labs no longer match. Three implications stand out:

The open-weight frontier keeps compressing. A 770B-parameter, 1M-context model with FP8 weights, vLLM day-one support, and sub-$3 output pricing would have been unthinkable as an open release a year ago. The gap between “best open” and “best closed” is now measured in months, not years.

MoE sparsity is the standard, not the exception. At 49B active parameters, Hy4 preview delivers flagship capacity at mid-tier serving cost — the same architecture bet made by DeepSeek, Qwen, and GLM. The 6.4% activation ratio is among the leanest seen at this scale.

Self-improvement claims are moving from hype to measurable artifact. A 31.8% throughput gain from model-driven optimization is a concrete, falsifiable number. If the industry’s trajectory holds, expect “did the model help build itself?” to become a standard question in release notes — and a genuine competitive moat for labs that can make it work.

For developers, the practical takeaway is simple: the FP8 weights are on Hugging Face today, vLLM 0.29.0+ can serve them, and the two-week free window on CodeBuddy is a low-cost way to kick the tires on a model Tencent bets can out-code its domestic rivals.

Sources