← All posts / Models

Tencent Open-Sources Hy4 Preview: A 770B MoE Flagship With 1M-Token Context

Tencent's Hunyuan team releases Hy4 preview, a 770B-parameter open-weight MoE with 49B active parameters, a 1M-token context window, and coding results that rival GLM-5.3 and Kimi K3.

Tencent Open-Sources Hy4 Preview: A 770B MoE Flagship With 1M-Token Context

On August 28, 2026, Tencent’s Hunyuan team quietly dropped one of the most consequential open-weight releases of the year: Hy4 preview, a next-generation flagship language model with 770 billion total parameters built on a Mixture-of-Experts (MoE) architecture. Unlike the tightly-held frontier models from Western labs, Hy4 preview ships with downloadable weights — and its benchmark profile suggests the gap between “open” and “closed” AI has narrowed again.

What Hy4 Preview Actually Is

Hy4 preview is a new-generation Mixture-of-Experts flagship developed by the Tencent Hy team. The headline numbers: 770B total parameters, with roughly 49B activated per token, and a context window that stretches to one million tokens. That active-parameter ratio — under 7% of the full model — is the whole point of modern MoE design: you get the knowledge capacity of a dense 770B model while paying the inference cost of something far smaller.

The model is positioned squarely at software engineering, research, and analytical workloads rather than casual chat. Tencent’s own announcement emphasizes agentic productivity: tool use, long-horizon reasoning, and the ability to work across entire codebases — which is exactly where a 1M-token context window stops being a spec-sheet flex and becomes practically useful. You can fit a large monorepo, its documentation, and its issue history into a single prompt and still have room for the actual task.

The release is available now on Hugging Face (tencent/Hy4-preview) and ModelScope (Tencent-Hunyuan/Hy4-preview), with vLLM recipes already published for production deployment. Community tracking pages picked up 40+ benchmark results within days of release.

Benchmark Performance

The numbers that got attention are the coding and engineering results:

  • SWE-Bench Pro (Public): 65.7% — this places Hy4 preview as the top-ranked open-source model on SWE-Bench Pro (rank #5 overall on LLM Stats’ leaderboard, score 0.657), ahead of every other open-weight competitor
  • HLE (Humanity’s Last Exam): 55.40 — a strong showing on one of the hardest reasoning evaluations still standing
  • GPQA Diamond results solidifying its science-reasoning credentials
  • In blind engineering evaluations, Hy4 preview scored an average of 2.99 out of 4.00, slightly ahead of GLM-5.3 (2.92) and Kimi K3 (2.94) — the two Chinese open-weight models that had been trading the open-source crown for months
  • On public aggregation leaderboards like BenchAlign, it debuted at #7 out of 228 models with 79.16/100

The context matters here. GLM-5.3 and Kimi K3 have spent 2026 leapfrogging each other at the top of the open-weight charts. Hy4 preview didn’t just join that race — it nudged past both in blind engineering tests on its first attempt. Independent comparisons against DeepSeek, Qwen 3.8 Max, GPT 5.6 Sol, and Claude Opus 5 across nearly thirty benchmarks show a model that is competitive with the global frontier on coding-adjacent tasks while remaining freely downloadable.

Why This Release Matters

Three reasons this is more than a routine model drop.

First, the open/closed gap keeps closing. A year ago, “open-weight” reliably meant “acceptably worse.” Today, the top open-source model on SWE-Bench Pro — one of the hardest coding benchmarks, built from real GitHub issues in production repositories with multi-file diffs — is a Chinese open-weight release. Frontier labs still lead on the absolute ceiling, but the practical difference for most engineering work has collapsed to noise.

Second, MoE efficiency at scale is proven out. Running 49B active parameters out of 770B total means serving costs per token are dramatically lower than a dense model of equivalent capability. Combined with vLLM support at launch, Hy4 preview is deployable on hardware budgets that dense frontier models simply don’t fit. Reports suggest the model even contributed to optimizing its own training and inference pipeline — recursive self-improvement in a narrow, engineering sense.

Third, the Chinese open-weight ecosystem is compounding. DeepSeek, Qwen, GLM, Kimi, and now Hunyuan have turned open-source releases into a steady cadence. Each release raises the floor for everyone — researchers get weights to study, startups get models to build on, and the releases pressure closed labs on pricing. Hy4 preview entering the ring with a win over GLM-5.3 and Kimi K3 guarantees the next round of one-upmanship is already underway.

The Caveats

It is called a preview for a reason. Early adopters note multimodal support has limits compared to native-multimodal competitors, and “preview” typically means the final version may differ — sometimes significantly — from what’s on Hugging Face today. Independent verification of benchmark claims is still landing; some leaderboard entries carry “estimated” evidence status pending reproduction. And serving a 770B model, even with sparse activation, is not trivial — realistic GPU requirements remain substantial for teams wanting to self-host at production latency.

None of these caveats change the core fact: a 770B-parameter, 1M-context, coding-first flagship just became free to download.

What to Watch

The obvious next milestones: independent eval reproductions over the coming weeks, the stable Hy4 release (presumably with fuller multimodality), and how GLM and Kimi respond — both have shown they iterate fast. If Hy4’s final version holds or extends this benchmark position, Tencent moves from cloud-AI player to genuine open-weight heavyweight in a single release cycle.

For engineers and researchers, the practical advice is simple: the weights are on Hugging Face, the vLLM recipe is published, and the SWE-Bench Pro numbers say this is currently the strongest open model for real repository-scale work. That combination doesn’t come along every week.