← All posts / Models

Six Models, Zero Secrets: Inside K2 Horizon, the Largest Fully Open AI Release in History

MBZUAI's Institute of Foundation Models shipped six Apache 2.0 models from 0.9B to 375B parameters — with weights, code, training data, and methodology all public. It's the widest disclosure ever from a frontier-adjacent lab.

Six Models, Zero Secrets: Inside K2 Horizon, the Largest Fully Open AI Release in History

Six Models, Zero Secrets: Inside K2 Horizon, the Largest Fully Open AI Release in History

On September 3, 2026, the Institute of Foundation Models (IFM) at Abu Dhabi’s Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) released K2 Horizon — a fleet of six AI foundation models spanning 0.9 billion to 375 billion parameters. That alone would be a routine week in an industry that ships frontier models like clockwork. What makes this launch historically unusual is what ships with the models: weights, training code, training data and recipes, configurations, evaluation results, and methodology — all of it public, all of it under Apache 2.0.

IFM calls it “the largest fully open-source model launch in AI history,” and the claim is hard to argue with. While most releases described as “open” stop at publishing weights — a download link and a license file — K2 Horizon opens the entire development trail. Reuters independently reported the release as an unusually comprehensive open-source effort focused on reproducibility rather than weight access alone.

Why this is different from “open weights”

The timing is pointed. The release lands in the middle of an industry debate about what “open” actually means, and IFM’s launch materials take an explicit swing at the “open weights dialogue that has dominated AI industry headlines over recent weeks.”

The distinction matters. An open-weights release lets you download a model and run it, but you cannot see how it was built. You cannot reproduce the training run, audit the data mixture, or study the intermediate checkpoints where capabilities emerged. You are trusting the lab’s benchmarks and its characterizations of the data. A fully open release inverts that: the claim becomes checkable.

“Open source is much more than open weights. Science works when others can see the data, follow the method, reproduce the result, and improve on it,” said Eric Xing, Founder of IFM and President of MBZUAI. “K2 Horizon delivers on that need. Every model in the fleet ships with its training data, recipe and evaluations. This is open science, and we believe it’s the best path forward for AI.”

The fleet: from a wristwatch to the data center

The six models share a core architecture, vocabulary, training methodology, interfaces, and deployment tooling (with a smaller vocabulary for the 0.9B variant). The idea is a continuous path from prototype to production without changing model family or workflows:

  • 0.9B — the industry’s best-performing model of its size in math, reasoning, and tool use, per IFM; small enough to run on a smartwatch. Aimed at edge deployment in energy, logistics, healthcare, and public services.
  • 3.7B — a fine-tuning-friendly size for developers, with reasoning that IFM calls the industry’s best under 4B, matching or beating many larger models.
  • 7B — claimed as the best-performing model under 10B parameters, with strong software-engineering and deep-research capability; runs on a phone.
  • 32B — a dense model built for heavy use on laptops and on-premise servers; among the most capable dense models for local hosting.
  • 36B-A4B — a sparse model activating only 4B parameters, introducing a new “Mixture of Value Attention” (MoVA) architecture that IFM says outperforms many larger models.
  • 375B-A23B — the flagship: a sparse Mixture-of-Experts model with 375B stored parameters, roughly 23B active per token, and a native 524,288-token context window, built to compete with leading open-weight models on reasoning and agentic work.

Per IFM, the three smallest models (0.9B, 3.7B, 7B) set new state of the art in their respective size classes across mathematics, reasoning, software engineering, and agentic capabilities. Community benchmark posts around the 375B flagship cite GPQA Diamond at 87% and Terminal-Bench 2.1 around 72% — strong numbers, though vendor-reported until independent evaluations accumulate.

Three new techniques under the hood

K2 Horizon’s technical story rests on three innovations IFM claims:

  1. Diffusion distillation — generates blocks of tokens in parallel rather than strictly one at a time, speeding inference by roughly 3X without degrading response quality.
  2. Mixture of Value Attention (MoVA) — an architectural change that improves reasoning without adding compute, powering the 36B-A4B model.
  3. Dynamic model routing — directs tasks to the most cost-effective model in the fleet automatically, giving developers a practical path from prototype to production.

The lineage is visible. This trajectory began with the original K2, continued through K2-Think and K2 Think V2, and reaches its widest expression with Horizon. Each generation opened more of the stack; Horizon opens essentially all of it.

The fine print on “fully open”

Honesty requires noting the timing nuances. IFM’s launch materials present the training lifecycle as part of the release, but the current Hugging Face card for the 375B-A23B flagship states the final checkpoint is available now and that intermediate checkpoints, data, and training code “will be released” — some artifacts are still being staged. Where redistribution rights prevent direct publication of a dataset, IFM says it provides source descriptions, filtering procedures, construction methods, and mixture composition instead of raw data.

And openness is not the same as cheapness. IFM’s own published SGLang serving recipe for the flagship is validated on an eight-H200 node with tensor and expert parallelism, BF16 precision, and FlashAttention-3. The weights may be free to download; production-grade inference remains a data-center-scale commitment.

Availability is broad: Hugging Face for weights, day-zero support in vLLM, SGLang, and Ollama, deployment paths on NVIDIA, AMD, and Cerebras hardware, and hosted APIs through inference partners including Compass, Cerebras, AWS, and Nebius.

Why a university lab is doing this

The strategic backdrop is the UAE’s national AI bet. IFM was established in May 2025 and now operates from three offices — Abu Dhabi, Silicon Valley, and Paris — placing frontier model development in the Emirates while pulling talent from established centers. Its portfolio includes the K2 series, the Arabic-English Jais models, and PAN, a world model for embodied reasoning.

“The most important technology of our time should be built with the world, not kept from it,” Xing said. “K2 Horizon is what that conviction looks like when an institution commits to it completely, and it reflects the environment we work in here in the UAE: a country that decided to build AI rather than wait for it.”

What to watch

The open-ecosystem race of 2026 is crowded — Kimi K3, GLM-5.2, Qwen 3.8, GPT-OSS, Nemotron 3 — but almost all of those are open-weight releases. K2 Horizon changes the terms of competition: if its numbers hold up under independent scrutiny, the question stops being “which open model is best?” and becomes “why is your lab still hiding the training data?”

The markers to watch: whether the promised intermediate checkpoints and data land on schedule; whether independent evaluations confirm the small-model SOTA claims, which are the most falsifiable; and whether dynamic routing actually works as a fleet-level product rather than a paper feature. If those checks pass, September 3, 2026 may be remembered as the day “open source AI” finally meant what it says.