← All posts / Models

29B Parameters, 4B Active, Zero NVIDIA: China Telecom Open-Sources Xing4.0, an Agent Model Trained Entirely on Ascend

China Telecom's Xing4.0-29B-A4B is the first ~30B-class open model trained end-to-end on Huawei Ascend 910C with MindSpore — and its agent benchmarks beat both Gemma4 and Qwen3.6 on agentic coding.

29B Parameters, 4B Active, Zero NVIDIA: China Telecom Open-Sources Xing4.0, an Agent Model Trained Entirely on Ascend

On September 17, 2026, China Telecom Artificial Intelligence Technology Co., Ltd. quietly uploaded a set of model weights to Hugging Face, ModelScope, and Modelers. No keynote, no livestream. But the artifact itself speaks loudly: Xing4.0-29B-A4B, a 29-billion-parameter Mixture-of-Experts language model that activates only 4B parameters per token, and — the part that matters far beyond benchmark tables — the first model of its scale trained entirely on Huawei’s Ascend NPU platform using the MindSpore framework, rather than on NVIDIA GPUs with CUDA software.

For anyone tracking the slow-motion decoupling of the US and Chinese AI stacks, this release is a data point worth reading carefully. It is not a frontier lab demonstrating a compute megaproject; it is a state-owned telecom operator shipping a mid-sized, aggressively open-sourced agent model built on domestic silicon, with benchmark numbers that genuinely compete in its weight class.

What was actually released

The Xing (星辰, “Stars”) series is the descendant of China Telecom’s TeleChat line, which had already made history in January 2026 as the first MoE models trained entirely on domestically designed chips. Xing4.0-29B-A4B is the newest generation, and it shipped in three variants simultaneously:

  • Xing4.0-29B-A4B — full weights under Apache 2.0 (~58.1 GB safetensors, max shard 4.1 GB)
  • Xing4.0-29B-A4B-FP8 — quantized training/inference variant
  • Xing4.0-29B-A4B-GGUF — consumer-local format; a Q8_0 build comes in around 12 GB, small enough to run on a single high-end consumer GPU

The architecture follows the now-standard efficient-agent recipe: 40 layers, 3,584 hidden size, MLA (Multi-head Latent Attention), 64 routed experts with 4 active per token plus 1 shared expert. The context story is aggressive for this class: 256K tokens natively, extensible to 512K, reflecting the reality that agent workloads — repository-scale coding, long tool-call chains, document grounding — live and die by context.

An architecture built for agents, not chatbots

Two design choices stand out in the model card. First, the team describes an “mHC + MLA + MTP” stack: Multi-head latent Attention for cheap long-context KV caches, MTP (Multi-Token Prediction) for speculative decoding at serve time, and mHC — a hybrid attention variant the team implemented as custom fused operators written in Ascend C, Huawei’s native kernel language. That last detail is easy to skim past, but it is the whole story of this release: features like mHC do not exist as ready-made kernels on the Ascend software stack the way they do in CUDA-land. China Telecom’s engineers had to build them, numerically align them across frameworks, and fuse them by hand.

Second, the model was explicitly tuned for “complex engineering tasks” — multi-step planning, tool calling, and long reasoning-chain execution — and the team did format-level adaptation for agent frameworks including OpenCode, Claude Code, OpenClaw, and Hermes. In other words, the intended deployment is not a chat window; it is an autonomous coding or operations agent running inside an existing harness.

The numbers: genuinely competitive in-class

China Telecom benchmarked Xing4.0-29B-A4B against the two obvious open-weight rivals in the efficient-MoE class, Google’s Gemma4-26B-A4B and Alibaba’s Qwen3.6-35B-A3B. The agentic results are the interesting ones:

BenchmarkXing4.0-29B-A4BGemma4-26B-A4BQwen3.6-35B-A3B
SWE-bench Verified75.053.076.0
SWE-bench Multilingual66.051.067.2
Terminal-Bench 2.157.530.051.5
Claw-Eval76.5571.4974.54
DeepresearchBII60.839.359.7
Tau3-Bench64.6358.967.2
AIME202690.088.392.7
IFBench69.6772.6765.5

The pattern is consistent: on agentic coding and tool-use evaluations — Terminal-Bench, Claw-Eval, DeepresearchBII — Xing4.0 takes the lead outright, and on SWE-bench Verified it sits one point behind Qwen3.6 while crushing Gemma4 by 22 points. Gemma4’s lead on IFBench (instruction-following) is the one clear deficit. Nobody would claim this is a frontier-class model, but “an Apache-2.0 29B MoE that hangs with Qwen3.6 on SWE-bench” is a legitimately strong 2026 release — and the fact that it was trained without touching a single NVIDIA GPU is the asterisk that makes it geopolitically significant.

The 96% number: what co-design actually buys

The most technically revealing claim in the release concerns training efficiency. Through what the team describes as multi-level co-optimization — fine-grained MoE communication optimization, selective recomputation, “DVM” automatic graph-operator fusion, and those hand-written Ascend C fused kernels for mHC — overall training throughput improved by roughly 96% over out-of-the-box Ascend performance. Read that as: nearly doubling throughput, not from better chips, but from software co-design.

This matters because the standard critique of non-NVIDIA training stacks is that raw peak FLOPS understate the real gap; kernel maturity, collective-communication tuning, and framework overhead widen it. A documented ~2x software-side gain on Ascend 910C clusters is exactly the kind of grind-it-out evidence that the domestic stack is closing the usability gap from below. It echoes the trajectory of iFlytek’s Spark X2-Flash and the DeepSeek-on-Ascend continuation-training work reported earlier in 2026 — different teams, same lesson: co-design partially substitutes for silicon advantage.

Ecosystem play: PRs everywhere, walled garden nowhere

The release also shows a maturing open-source strategy. Xing4.0 support has been submitted (pending merge at release time) to SGLang (#39793), vLLM (#57135), TensorRT-LLM (#19283), llama.cpp (#29012), and KTransformers (#2168) — with ready-to-paste serve commands including OpenAI-compatible APIs, MTP/EAGLE speculative decoding, and a xing4 reasoning/tool-call parser. Fine-tuning paths cover both LLaMA-Factory (the global default) and MindFormers (the Ascend-native path), and multi-chip deployment is handled via BAAI’s FlagOS for cross-architecture one-click rollout. There is even a KTransformers recipe for running the model on GPU/CPU heterogeneous setups with just 24 GPU-resident experts.

For Chinese enterprises under procurement pressure to de-NVIDIA, this is the pitch: a carrier-grade company standing behind an Apache-2.0 model whose entire lineage — chips, framework, operators, weights — is domestic, deployable on rented state-cloud Ascend clusters or on whatever hardware survives the latest export-control whiplash.

Why it matters

Strip away the nationalism and the engineering story is still notable: the efficient open-model frontier (small-total, small-active MoE with long context) has a new entrant that wins real agent benchmarks per parameter. Add the geopolitics back in and the release becomes a marker: the “trained entirely on domestic compute” claim has now moved from trillion-parameter bragging rights down to the 30B workhorse class — the tier that actually gets deployed in volume by banks, telecoms, and government systems integrators.

The open question is iteration speed. NVIDIA’s moat was never a single training run; it was the compounding of thousands of validated kernels, toolchains, and operator habits. One excellent 29B model on Ascend proves feasibility. A cadence of them — Xing4.1, Xing5, on time, with this benchmark discipline — would prove a stack. China Telecom, of all companies, understands the value of network effects better than most.