← All posts / Models

Tencent Open-Sources Hy4 Preview: A 770B MoE Flagship Built for Real Work

Tencent's Hunyuan team releases Hy4 preview, a 770B-parameter MoE model with 49B active parameters, 1M-token context and an early recursive self-improvement loop.

Tencent Open-Sources Hy4 Preview: A 770B MoE Flagship Built for Real Work

Just one day before the end of August 2026, China’s open-weight AI race produced its largest salvo yet. On August 28, Tencent officially released and open-sourced Hy4 preview, the next-generation flagship large language model from its Hunyuan team — a 770-billion-parameter Mixture-of-Experts system with 49B parameters activated per token and a context window exceeding one million tokens.

The release is more than another big number. It is Tencent’s clearest statement yet that it intends to compete at the frontier of open-source models, and it arrives with an unusual twist: Hy4 preview played a direct role in its own development, participating in the automated optimization of its training methods, data strategies, evaluation frameworks, and low-level operators.

What Tencent shipped

Hy4 preview is a Mixture-of-Experts (MoE) architecture, following the same broad design philosophy as DeepSeek’s flagship models: enormous total parameter count, aggressive sparsity, and a long context window. According to the vLLM recipe page for the model, the backbone spans 78 layers, with the first layer using a dense feed-forward network and the remaining 77 layers using sparse expert routing.

The headline specifications:

  • 770B total parameters, 49B activated per token
  • 1M+ token context window
  • Released open-weight under Apache License 2.0 on Hugging Face, with an FP8 quantized variant (Hy4 preview-FP8) published alongside the base weights
  • Available globally via Tencent Cloud TokenHub and OpenRouter, plus Tencent’s own products: WorkBuddy, CodeBuddy, Yuanbao, and ima
  • Free to use on WorkBuddy and CodeBuddy for two weeks after launch, with free access to the previous Hy3 generation extended to September 30

API pricing undercuts most frontier competitors: USD 0.834 per million input tokens, USD 2.501 per million output tokens, and USD 0.042 per million tokens for cache hits. That cache-hit price in particular is aggressive enough that high-volume agentic workloads — which re-send large prompt prefixes on every tool call — will find it difficult to ignore.

Benchmarks: top tier of open weight, with caveats

Tencent’s announcement is measured in its claims, saying the model ranks “among the top tier of open-source models” rather than claiming an outright crown. The numbers that have surfaced so far are consistent with that positioning.

In an internal blind evaluation conducted by Tencent — 163 experts assessing 203 engineering tasks — Hy4 preview scored an average of 2.99 out of 4.00, slightly ahead of GLM-5.3 (2.92) and Kimi K3 (2.94). That is a narrow margin, and it is Tencent’s own evaluation, but the choice to name competitors explicitly is itself a signal of confidence.

On public leaderboards, third-party tracking shows Hy4 preview at roughly 65.7% on SWE-bench Pro (public), and community-reported figures circulating around the launch cited scores in the region of 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-bench Multilingual. On the BenchAlign aggregate leaderboard it was ranked around #7 of 228 models within a day of release. Independent verification is still catching up — the model is less than a day old at time of writing — but the early picture is of a model that genuinely competes with the strongest open-weight systems from DeepSeek, Alibaba’s Qwen, and Moonshot, and presses into the lower edge of the Western closed frontier.

The stated focus is “real productivity scenarios” rather than leaderboard peacocking: software engineering, office analysis, game development, and scientific research.

The recursive self-improvement loop

The most technically interesting claim in the release is buried mid-announcement: Hy4 preview contributed to its own development process. Tencent says the model participated for the first time in the automated optimization of training methods, data strategies, evaluation frameworks, and low-level operators — proposing approaches, running experiments, and iterating on results, with the resulting code, logs, and feedback feeding into subsequent rounds of exploration. Tencent describes this as “an early-stage recursive self-improvement loop.”

The model also autonomously analyzed bottlenecks in its inference system and ran multiple optimization rounds on operator fusion and communication patterns. The result: a 31.8% increase in end-to-end throughput over the baseline, with gains holding across different context lengths and concurrency levels.

This is the kind of claim that warrants skepticism until independently replicated — “the model improved its own training” is easy to overstate. But even a partially automated pipeline in which the model proposes and tests improvements to its own training stack is a meaningful organizational achievement, and it echoes a broader 2026 trend of AI accelerating AI research itself.

Built by the Yao Shunyu era of Hunyuan

Hy4 preview is the latest step in Tencent’s accelerated model push since it hired Shunyu Yao, the former OpenAI researcher best known as a co-author on ReAct, SWE-agent, and the Swarm agent framework, to lead its foundation model efforts. The Hunyuan 3.0 release in April was the first major fruit of that hire; Hy3 preview followed within weeks; and the cadence has not let up since. The Hy4 series is expected to roll out additional models soon under a preview-first, then official-release strategy that feeds real-world usage back into training.

The training data itself was “co-created” with Tencent domain experts across software engineering, gaming, finance, and security — a vertical-data strategy aimed at tasks that general web crawls handle poorly, such as debugging a long-context codebase or producing a coherent financial analysis across scattered documents.

Why it matters

Three implications stand out.

First, the open-weight frontier keeps closing the gap. A 770B Apache-2.0 model with 1M context, agentic coding strength, and sub-dollar input pricing would have been unthinkable from a Chinese lab openly published two years ago. Enterprises that need frontier-adjacent capability without sending data to a US API now have another serious option.

Second, China’s internal competition is now the main event. Hy4 preview’s most direct rivals are GLM-5.3 (Zhipu/Z.ai), Kimi K3 (Moonshot), DeepSeek’s flagships, and Qwen3.8 (Alibaba) — all Chinese, all open or hybrid. Tencent’s explicit head-to-head blind eval against GLM-5.3 and Kimi K3 shows the domestic race is being fought in the open.

Third, self-improvement is moving from paper to production. The 31.8% inference-throughput gain from the model optimizing its own serving stack is exactly the kind of concrete, measurable result that separates marketing from substance. If that loop generalizes, it changes the economics of running frontier-scale models — and it was the model, not just the engineers, that found the wins.

Hy4 preview is available now on Hugging Face (tencent/Hy4-preview), with the FP8 variant for deployments that need it. Whether it dethrones DeepSeek at the top of the open-weight pile will take a few weeks of independent benchmarks to settle — but August 2026 will be remembered as the month the open-source frontier hit 770B.