← All posts / Tools

75GB of RAM Later: Microsoft Puts a 137B-Parameter Coding Agent on Your Laptop

At its Windows and Surface event, Microsoft announced that GitHub Copilot will soon route tasks between cloud and a local, quantized MAI Code 1.1 Flash — a 137B-parameter MoE coding model that fits in 53GB — sandboxed by the new open-source Microsoft Execution Containers.

75GB of RAM Later: Microsoft Puts a 137B-Parameter Coding Agent on Your Laptop

At its Windows and Surface event in San Francisco on October 7, Microsoft made the case that the next chapter of the PC is local AI. Satya Nadella shared the stage with NVIDIA’s Jensen Huang to show off “hybrid intelligence” — and the most consequential developer-facing announcement wasn’t a new laptop. It was a plan to run a 137-billion-parameter coding model entirely on a device, and to cage the agents that use it with a new open-source sandboxing library.

Coming by the end of October, GitHub Copilot will decide for itself when a task is best handled by on-device intelligence and when it should call cloud-scale models. Behind that promise sit two pieces of engineering: a heavily optimized local build of MAI Code 1.1 Flash, and Microsoft Execution Containers (MXC), a policy-driven containment system for agentic coding sessions.

The model: 137B parameters, 6.8B active, 53GB on disk

MAI Code 1.1 Flash is a coding-optimized mixture-of-experts model from Microsoft AI with 137 billion total parameters but only 6.8 billion active per token. The cloud variant in BFloat16 weighs in at roughly 265GB; the on-device version Microsoft is shipping uses mixed-precision quantization at approximately 3.3 bits per weight, plus DFlash2 sliding-window speculative decoding to raise decode throughput.

The result is a 53GB footprint — an 80% size reduction — that still hits hard numbers on agentic benchmarks. On Surface Laptop Ultra (the new NVIDIA RTX Spark machine with up to 128GB of unified memory and up to 1 petaflop of AI compute), Microsoft reports:

  • SWE-Bench Verified: 70.8% on-device vs. 72.6% for the full-precision cloud model — and versus just 32.0% for Unsloth’s GPT-OSS-120B GGUF, the popular local alternative
  • Terminal-Bench 2.1: 66.29% on-device vs. 62.9% in the cloud and 23.6% for GPT-OSS-120B
  • Throughput: 923.5 tokens/second prompt processing at 64k context, 769.8 at 128k
  • Peak memory: 75.5GB at 256k context — the KV cache and runtime share that unified pool with the OS and your apps

The performance framing matters. Microsoft is explicit that quantization for code is not a graceful-degradation problem: one wrong token means a syntax error, a malformed tool call, or a broken diff. So the evaluation measures task completion, not perplexity. And on those terms, the quantized model doesn’t just survive — on Terminal-Bench it actually beats the BFloat16 cloud variant, presumably a side effect of the evaluation setup and decoding strategy rather than quantization magic.

There is a hardware asterisk, and Microsoft doesn’t hide it: for best performance the company recommends devices with more than 120GB of RAM. Today that effectively means the RTX Spark-based Surface Laptop Ultra and the new Surface RTX Spark Dev Box and Work Station. Local frontier coding is arriving, but it’s arriving on machines that cost as much as a used car.

The sandbox: Microsoft Execution Containers

Local inference alone doesn’t make a session safe. An agent’s shell commands inherit the full access of the account running them — moving the model onto the device changes nothing about that. So Microsoft is pairing local models with MXC, an open-source library from the Windows team that translates a declarative policy into native OS controls.

The design is deliberately boring in the best way:

  • On Windows, GitHub Copilot uses the BaseContainer tier of MXC’s ProcessContainer backend
  • On macOS, it uses Seatbelt — Apple’s long-standing sandbox
  • On Linux, it uses bubblewrap

No separate VM or container image is required; the backends are OS-native process boundaries. When sandboxing is enabled, shell commands and — by default — local MCP servers and language servers run inside the boundary. Built-in file tools are checked in-process by the Copilot harness, and remote MCP servers remain outside the local sandbox with connection-policy checks applied in-process. A /sandbox slash command in the Copilot CLI exposes the configuration at any time.

This is the piece with legs beyond Microsoft’s hardware. MXC is open source, policy-driven, and cross-platform in its backend selection. As agentic coding becomes the default mode of development, “which OS primitives contain my agent” becomes a serious question, and Microsoft just shipped a reference answer.

Hybrid orchestration, not offline mode

The second Copilot feature is routing intelligence. Developers get two modes: Auto orchestration, where Copilot decides task-by-task whether to use local or cloud inference (preserving cached work across the session as it shifts), and explicit local-model selection through the Windows ML provider or any OpenAI-compatible local endpoint.

Microsoft is careful about the boundaries: local inference does not make the session offline, model selection is independent of tool execution, and the sandbox policy applies regardless of which model requested the work. The demo — a daily automation that reads repositories, runs tests in a sandboxed working copy with no network access, and writes an HTML triage dashboard — ran entirely on-device with MAI Code 1.1 Flash selected as the model.

Why it matters

Three currents converge here. First, the edge-inference stack is maturing fast enough that a 137B MoE quantized to 3.3 bits is a practical daily-driver coding model rather than a demo. Second, agent security is finally getting OS-level treatment instead of prompt-level hoping — MXC’s appearance at a flagship consumer event, alongside an announcement that Nadella framed around “building Windows for hybrid intelligence,” signals that containment is becoming a platform feature rather than an enterprise add-on. Third, the economics: MAI Code 1.1 Flash was already billed as “a quarter of the cost” of cloud peers, and local execution moves that cost to a fixed hardware investment.

The caveat is hardware gating. A model that needs 120GB+ of RAM to shine is, for now, a halo feature for expensive machines — the same playbook as Apple’s local-Apple-Intelligence rollouts, but aimed at developers. Still, the direction is unmistakable: the boundary between “cloud agent” and “local agent” is dissolving into an orchestration decision, and Microsoft is betting it owns both ends.

Local models and sandboxed tools are rolling out to GitHub Copilot CLI, the Copilot app, and VS Code now, with the auto local/cloud routing arriving by the end of the month.