← All posts / Models

Saudi Arabia's HUMAIN Bets on China: humain-m3 Is a 428B-Parameter Arabic Frontier Built on MiniMax

HUMAIN's humain-m3 — a 428B-parameter Arabic MoE adapted from China's MiniMax-M3 — averages 89.37% across seven Arabic benchmarks, beating GPT-5.6 SOL and Opus 5 in the company's own tests.

Saudi Arabia's HUMAIN Bets on China: humain-m3 Is a 428B-Parameter Arabic Frontier Built on MiniMax

RIYADH — The most consequential model launch at LEAP 2026 was not a new flagship from a Western lab. It was a 428-billion-parameter mixture-of-experts system, commissioned by Saudi Arabia’s state-backed AI company HUMAIN and delivered by China’s MiniMax, that now claims the top average score on Arabic-language AI evaluations — ahead of GPT-5.6 SOL and Anthropic’s Opus 5 in the company’s own benchmark table.

The model, humain-m3, went into research preview on HUMAIN Node on Thursday, September 3. It is the clearest signal yet of where sovereign AI is actually heading: not every nation will train frontier models from scratch, and the fastest path to frontier-grade language capability may run through a Chinese foundation model, one trillion tokens of Arabic-native pre-training, and a sovereign inference platform on top.

What HUMAIN actually built

humain-m3 is a 428-billion-parameter mixture-of-experts (MoE) model built on the MiniMax-M3 lineage. The architecture holds hundreds of billions of parameters in reserve, but activates only about 23 billion per token — the efficiency trick that has made MoE the dominant design at the frontier, keeping inference costs manageable while preserving total capacity.

The base model is MiniMax-M3, released by the Shanghai lab in June 2026 as a native multimodal system with a 1M-token context window, ~428B total parameters, and ~23B activated. HUMAIN’s contribution was not a new base architecture. It commissioned MiniMax to adapt M3 for Arabic: further pre-training on more than one trillion tokens of native Arabic content, followed by the alignment and safety work needed to serve Kingdom-based workloads.

That distinction matters. HUMAIN is explicitly calling this a model “built on the MiniMax-M3 lineage” — a commissioned adaptation, not a from-scratch national LLM. It joins the company’s existing ALLAM family of Arabic models, which includes the ALLAM 34B that ranked #2 on the BALSAM Arabic leaderboard earlier this year, right behind GPT-5.2.

The benchmark claim: 89.37% across seven Arabic tests

In HUMAIN’s own evaluation, humain-m3 averaged 89.37% across seven equally weighted public Arabic benchmarks, compared with 80.34% for the MiniMax-M3 reference checkpoint — a 9.03-point lift attributable to the Arabic-native pre-training. The company’s table also placed it above GPT-5.6 SOL at 87.30% and Opus 5 at 87.34%, and said the model led five of the seven individual tests.

The benchmark suite spans Arabic understanding, native and translated knowledge, academic exams, language proficiency, truthfulness, and retrieval-augmented generation — a broad spread, though the aggregate is still one number. Two caveats deserve emphasis: these are preview-checkpoint results from HUMAIN’s own evaluation, not an independent ranking, and a single average cannot substitute for testing on a specific deployment’s tasks. But the pattern — a Chinese frontier base, aggressively tuned for Arabic, outscoring the top Western closed models on Arabic work — is the story, regardless of decimal places.

Access today, weights tomorrow (maybe)

The practical product right now is access through HUMAIN Node, the company’s enterprise platform for model access and scaled inference. Developers get a no-code playground and an OpenAI-compatible API. Two tiers are live:

  • Research preview — the full checkpoint with thinking and streaming enabled, tool use and computer use for long-running agent workflows, three thinking modes (always-on, adaptive, off), and multimodal capabilities across text, image, and video, including long-video understanding and native screen operation.
  • Limited preview — adds a Saudi alignment guardrail, disables thinking and streaming, and runs with higher latency.

The more consequential commitment is the weight release. HUMAIN says it plans to publish downloadable weights under the MiniMax Community License, with a target of next month — but explicitly conditional on completing safety training and alignment work first. The preview exists partly to gather feedback on capability, safety, and alignment across Arabic dialects before general availability. Until those weights land, “open” remains a promise rather than a product.

Why a Saudi national model runs on Chinese foundations

The geopolitics is the subtext Bloomberg’s headline flagged: Saudi Arabia’s Public Investment Fund-backed national AI champion chose a Chinese foundation model for its flagship Arabic release, in the same week it signed megawatt-scale infrastructure deals with AMD, Cisco, AWS, and others.

The logic is pragmatic. Arabic is a lower-resource language for most Western frontier labs; the top closed models are strong but not Arabic-optimized. MiniMax-M3 ships as open weights under a community license, which makes commissioned adaptation legally straightforward and keeps a path to sovereign weight ownership. And HUMAIN’s stated strategy is a model-agnostic marketplace — architectures as interchangeable layers on top of Kingdom-controlled compute, data, and security. In that framing, picking the best available foundation for the job, regardless of origin, is not a compromise. It is the design.

It also continues a pattern: China’s open-weight ecosystem — Qwen, DeepSeek, Kimi, MiniMax — is becoming the substrate that other nations build sovereign AI on. When the weight release happens, humain-m3 will be one of the largest openly licensed Arabic-capable models in existence.

What to watch

Three things will determine whether this launch matters beyond the news cycle. First, the weight release: whether HUMAIN hits its “next month” target under the MiniMax Community License, and what the license actually permits. Second, independent verification: third-party Arabic evaluations to test the 89.37% claim against GPT-5.6 and Opus 5 on neutral ground. Third, adoption through HUMAIN Node — whether Saudi enterprises and government workloads actually standardize on an Arabic-adapted MoE, or keep defaulting to Western APIs.

For the model industry, the message is simpler. The frontier is fragmenting into regional frontiers, and the foundation layer just went global: a Gulf sovereign fund, a Shanghai AI lab, one trillion Arabic tokens, and a benchmark table where the familiar American names no longer sit on top.