Meta's Comeback Play: Muse Code, Muse Spark 1.2, and the Open-Weights Muse Glimmer
Meta Superintelligence Labs has shipped Muse Code (beta), a terminal coding agent powered by Muse Spark 1.2, and open-sourced the 30B Muse Glimmer under Apache 2.0 — the company's full return to open weights after its closed-API pivot.
Meta is making its most aggressive AI push since the Llama era — and this time it comes in three parts. Meta Superintelligence Labs (MSL), the division built around former Scale AI chief Alexandr Wang, has released Muse Code (beta), a terminal coding agent powered by the new Muse Spark 1.2 model, alongside Muse Glimmer, a 30-billion-parameter agentic model whose weights are now downloadable under a permissive Apache 2.0 license. Together, the three releases mark a coordinated attempt to reclaim the open-source relevance Meta surrendered when it shipped its first closed, API-only frontier model earlier this year.
What shipped, exactly
The centerpiece is Muse Code, announced August 5, 2026 as an early beta. It is a terminal-based coding agent — installed on macOS or Linux with a single curl command — built for what Meta calls “long-horizon software engineering”: planning changes, writing code, and validating the results across large repositories. Meta reported an 82.9 percent score on its own Terminal Bench coding test, putting it in direct competition with Claude Code, Codex, and the open-source coding agents that have proliferated this year.
Under the hood sits Muse Spark 1.2, a coding-focused update to Muse Spark 1.1 with a 1 million token context window. Meta says it significantly scaled up training compute on coding tasks, and — in an interesting twist — co-trained the model with the Muse Code harness itself, feeding rejection-sampled agent trajectories back into training so the model performs best when paired with its own agent runtime.
The third piece, Muse Glimmer, arrived August 10 and is arguably the most strategically significant. It is a 30B-parameter model purpose-built for always-on local agent workflows — small enough to run on a Mac or a PC with a single consumer GPU. Meta open-sourced the weights on Hugging Face under Apache 2.0, ungated, with optimized integrations for llama.cpp, MLX, and ExecuTorch landing within days. AMD announced day-one support on its Ryzen AI Max+ SoCs, including the new Ryzen AI Halo small-form-factor desktop.
The engineering that makes Glimmer run locally
The Muse Glimmer blog post is unusually candid about the constraints. A 30B model at full precision would need over 55 GB of memory — far beyond any consumer GPU. Meta solved this with quantization techniques that compress the language model to roughly 4-bit precision, shrinking it to under 20 GB while retaining a dedicated perception encoder for multimodal input. The result: interleaved text and image understanding, screenshot interpretation, and chart reading, all on hardware like an RTX 5080 or RX 7900 XTX with 24 GB of VRAM.
Glimmer was trained in three phases. Pre-training used logit distillation from Muse Spark — the larger teacher model — on a similar data mix. Mid-training shifted to longer-context, agent-heavy data with richer reasoning traces. Post-training combined supervised fine-tuning with on-policy distillation and reinforcement learning across general, reasoning, coding, and agentic domains. Meta evaluated it against Google’s Gemma4-31B and Alibaba’s Qwen3.6-27B, claiming strong results in its size class on agentic benchmarks including DeepSearch QA, MCP-Atlas, τ-Bench, and SWE-Bench.
The capability list reads like a checklist for local autonomy: end-to-end task completion, reliable tool calling with precise schemas, multi-step reasoning over long horizons, failure recovery (diagnosing and retrying failed tool calls rather than halting), scaffold compatibility with OpenClaw and other orchestration patterns, controllable reasoning effort, and training data spanning more than 100 languages.
Muse Code’s runtime: persistent subagents and a replayable event log
Muse Code’s architecture differs from the spawn-per-task pattern most coding agents use. It runs a simple main agent loop plus a set of async background agents that stay alive for the entire session, carrying out next steps and choosing when to report back. Meta argues this persistence cuts latency and reduces the need for human steering on difficult multi-step tasks.
Equally notable is the runtime design: every model call, tool run, approval, and edit is appended to a local event log that serves as a single source of truth. The runtime is replay-exact and restart-safe — after a crash, the agent resumes precisely where it stopped. That is a meaningful property for tasks that run for hours, and Meta demonstrated exactly that: in a kernel-optimization case study, Muse Spark 1.2 iteratively optimized GPU kernels over 1,000+ tool calls across up to 24 hours, writing, compiling, and profiling Triton implementations of KDA and MLA kernels for NVIDIA Hopper GPUs, with the agent re-deriving the algorithms from scratch rather than wrapping existing libraries.
Muse Code also ships with bundled skills — /plan turns a task into an approval-gated plan, /grill stress-tests the plan until it holds up, and /goal works toward a specified objective — plus support for spawning multiple persistent subagents per task.
Why this matters: Meta’s return to open weights
Context matters here. Meta built the modern open-weights movement on Llama, then stumbled: Llama 4 fell behind, Llama 4 Behemoth remains unreleased more than a year after announcement, and the company’s April pivot to a closed, API-only model with Muse Spark 1.0 disappointed the developer community that had optimized for Llama for years. Moor Insights analyst Anshel Sag called the new announcements “a return to form,” and CNBC framed the open-weights move as Zuckerberg pushing for U.S. leadership in open AI.
The timing is not accidental. Nearly every large open-weight release this year has come from a Chinese lab — DeepSeek, Alibaba’s Qwen, Moonshot, and Zhipu — while Western labs largely kept their best models closed. By open-sourcing a competitive 30B agentic model under Apache 2.0 (a more permissive license than Llama’s community license, with no gating), Meta is positioning itself as the Western counterweight, and Muse Spark 1.2’s weights are slated to follow under a modified Llama Community License.
There is also a commercial logic. Muse Spark 1.2 is available both inside Muse Code and through the Meta Model API, and Moor Insights notes its cost-per-task pricing is competitive. Coding agents have become the killer app for frontier models this year; Meta’s co-training approach — model and harness optimized for each other — mirrors the vertical integration that made Claude Code and Codex formidable.
The caveats
Temper enthusiasm with the usual flags. The 82.9 percent Terminal Bench figure is Meta’s own reporting on its own benchmark. Muse Code is a beta. Muse Spark 1.2’s open weights are promised but not yet shipped at the time of writing, and “modified Llama Community License” means restrictions may still apply compared with Glimmer’s clean Apache 2.0. And Muse Glimmer’s comparisons target same-size-class models — nobody is claiming it beats a frontier model, only that it brings frontier-adjacent agentic behavior to hardware people already own.
Still, the direction is clear. After more than a year of watching from the sidelines, Meta is shipping again — a coding agent, a coding model tuned for that agent, and a genuinely open local model, all within one week. The open-source AI landscape, dominated by Chinese labs since spring, just got a serious American entrant back in the game.
Sources
- [1] https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2
- [2] https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
- [3] https://www.cnbc.com/2026/08/10/meta-muse-glimmer-open-weight-ai.html
- [4] https://moorinsightsstrategy.com/field-notes/metas-muse-glimmer-and-recapturing-the-spark-of-llama-with-open-weights/