← All posts / Models

Meta Open-Sources Muse Glimmer: A 30B Agent Model You Can Run on One GPU

Meta's new Apache 2.0 Muse Glimmer model brings agentic AI to consumer hardware, paired with a 6,500-word Zuckerberg manifesto on personal superintelligence.

Meta Open-Sources Muse Glimmer: A 30B Agent Model You Can Run on One GPU

On August 10, 2026, Meta Platforms released Muse Glimmer, a 30-billion-parameter open-weight AI model designed specifically to run autonomous agents locally on consumer hardware. The model shipped under a permissive Apache 2.0 license with weights available for free download on Hugging Face, accompanied by a sweeping 6,500-word essay from CEO Mark Zuckerberg outlining his vision for “personal superintelligence for everyone.”

The release marks a notable strategic moment for the company. Just four months earlier, Meta had broken with its open-source tradition by releasing Muse Spark as a closed model — the first time in five years that a flagship Meta model couldn’t be freely downloaded. Muse Glimmer represents a return to open weights, even as Zuckerberg pledged that the larger Muse Spark 1.2 would also see an open release in the coming weeks.

What makes Muse Glimmer different

Muse Glimmer is not positioned as another general-purpose chatbot. Meta describes it as a dense, multimodal, agent-focused model — a 29.6-billion-parameter causal language model distilled from the larger Muse Spark, equipped with a dedicated perception encoder for vision understanding and a 131K-token context window. Crucially, it was built from the ground up for what Meta calls “always-on local agent workflows” — persistent, background-running AI agents that can use tools, browse the web, write code, and take actions on a user’s behalf.

The practical pitch is striking: Muse Glimmer fits on a single consumer GPU with 24 GB of VRAM. That means a developer or researcher with a high-end workstation can download the model, run it entirely offline, and build agent applications without any API calls, rate limits, or recurring costs. NVIDIA published a developer guide the same day detailing how to run agentic workflows with Muse Glimmer on its hardware.

The model includes an explicit reasoning mode — similar in spirit to the chain-of-thought approaches used by OpenAI’s o-series and Anthropic’s Claude — which Meta says can be toggled on for complex multi-step problems and off for simpler tasks, reducing unnecessary token consumption.

Benchmark performance

Meta released its own benchmark tables alongside the model, comparing Muse Glimmer against similarly-sized competitors. The results suggest strong agentic capabilities:

  • DeepSearch QA: 74.6% — outperforming the next-best open model (71.1%)
  • SWE-Bench Verified: 76.0% — a coding benchmark where models must resolve real GitHub issues
  • SWE-Bench Pro: 51.2% — a harder variant of the software engineering benchmark
  • τ³-Banking (Tau3-Bench): 24% — a tool-use benchmark where Glimmer beat Gemini 3.5 Flash-Lite (18%) and Qwen3.6-27B (16.7%)
  • SkillsBench: 44.3 — slightly behind Qwen3.6-27B’s 46.6
  • GAIA2, IFBench, AIME — Glimmer reportedly leads many reasoning and long-context rows

Independent reviewers noted that Glimmer performs particularly well on agentic tool-use tasks, which is exactly the niche Meta targeted. On pure knowledge and instruction-following benchmarks like SkillsBench, Qwen3.6-27B maintains a slight edge. Reddit’s r/LocalLLaMA community, a demanding audience for open-weight models, reported that the index-quality results met expectations while praising the roughly 20% reduction in reasoning tokens compared to comparable models.

It’s worth noting that most benchmark figures come from Meta’s own evaluation suite, and some tracking sites like benchlm.ai declined to publish an overall aggregate score pending independent verification.

Zuckerberg’s manifesto: personal superintelligence

The model release was paired with a substantial essay published on meta.com titled “The Future is for Everyone.” In approximately 6,500 words, Zuckerberg laid out an argument that superintelligent AI should not be concentrated in a handful of corporate labs or government-controlled systems. Instead, he championed the concept of “personal superintelligence” — AI agents that run locally, respect individual privacy, and work toward each user’s own goals rather than those of a platform.

Key themes from the essay include:

  • Decentralization over concentration: Zuckerberg argues that closed, centralized AI creates unacceptable power dynamics, and that open-weight models serve as a democratizing counterforce.
  • Privacy through locality: By running agents on personal devices rather than in the cloud, users gain both privacy and autonomy. Muse Glimmer’s single-GPU footprint is the technical embodiment of this principle.
  • American leadership in open AI: The essay frames open-source AI as a strategic asset for U.S. competitiveness, implicitly countering the narrative that open weights are primarily a safety risk.
  • Defense of model release: Zuckerberg directly addresses the safety debate, arguing that responsible open releases, coupled with robust evaluation, outweigh the risks of an AI oligopoly.

TechCrunch noted that the essay offers the clearest articulation yet of Zuckerberg’s “personal intelligence” vision — a future where every individual has a powerful, private AI agent as a kind of digital extension of themselves.

The competitive landscape

Muse Glimmer enters a crowded open-weight field. The 25B–35B parameter range is currently one of the most competitive tiers in open AI, with Alibaba’s Qwen3.6-27B and Google’s Gemma 4 serving as the primary rivals. Meta’s differentiation is the agent-first design: rather than optimizing purely for chat quality or knowledge recall, Glimmer was tuned for sustained, multi-step tool use.

This positioning aligns with the broader industry trend toward agentic AI. OpenAI, Anthropic, and Google have all invested heavily in agent capabilities in 2026, but those systems run in the cloud behind APIs. Meta’s bet is that a meaningful segment of developers — and eventually consumers — will prefer agents they can run and control locally.

Availability and what’s next

Muse Glimmer is available now on Hugging Face at meta-models/Muse-Glimmer-30B, with code also published on GitHub under Apache 2.0. NVIDIA has published integration guides for running the model on its RTX and data-center GPUs.

Meta indicated that the larger Muse Spark 1.2 — the model that made headlines in April for its closed release — would receive open weights “in the coming weeks.” If that holds, it would represent one of the most significant open-weight releases of the year and a strong signal that Meta intends to compete on openness as a core strategy, not just capability.

For the open-source community, Muse Glimmer is a welcome course correction. After the Muse Spark controversy, many observers questioned whether Meta had abandoned its open-weight commitments entirely. The Glimmer release — capable, genuinely open, and practically runnable — suggests those commitments remain alive, even as the debates over safety, regulation, and the responsible limits of openness continue.