← All posts / Tools

A 125B Model on a 12GB Gaming Card: How Strata Splits 24,576 Experts Across Your Whole PC

The open-source Strata engine runs Qwen3.8-Flash-Next, a 125B-parameter MoE model, on a single 12GB consumer GPU with 64GB of RAM at 60-95 tokens per second. Here is how expert offloading and speculative drafting make it work.

A 125B Model on a 12GB Gaming Card: How Strata Splits 24,576 Experts Across Your Whole PC

For years, the rule of thumb for local LLM inference has been brutal: one parameter costs roughly two bytes of GPU memory at 16-bit, so a 125-billion-parameter model needs somewhere north of 250GB of VRAM. That is server territory — multi-GPU nodes, HBM-stacked accelerators, and a power bill to match. This week, an open-source project called Strata landed on the Hacker News front page by breaking that rule in the most mundane way possible: it runs Qwen3.8-Flash-Next, a 125B-parameter mixture-of-experts model, on an ordinary gaming PC with a 12GB graphics card and 64GB of system RAM.

The numbers from the project’s own benchmarks are the kind that make local-inference enthusiasts do a double take. On an NVIDIA RTX 5070 (12GB) paired with a Ryzen 5 7600 and 64GB of RAM, Strata generates between 53 and 94 tokens per second depending on the quantization — comfortably faster than most people can read. Prompt ingestion, the metric that usually kills big local models when you paste in a long document, runs at 1,620 to 2,650 tokens per second on a 32K-token prompt. On the AMD side, an RX 9070 XT (16GB) with an older Ryzen 9 3900X still manages 44-60 tokens per second of generation.

The trick: stop treating VRAM as the only memory

The reason a 125B model normally needs a server is that dense models must keep every parameter resident on the accelerator. But Qwen3.8-Flash-Next is not dense — it is a MoE architecture with 24,576 small “expert” modules, of which only 10 fire per token. At any given moment, the overwhelming majority of the model is doing nothing.

Strata’s core insight is to treat the whole PC as a memory hierarchy rather than a GPU with a RAM afterthought. The project’s own explanation uses a kitchen metaphor: the things you use all the time stay on the counter, and the rest waits in the pantry. Concretely, three tiers cooperate:

  • The GPU (12-24GB) holds the few thousand experts that historically fire most often — the hot set.
  • System RAM (32-96GB) holds all 24,576 experts, with the CPU working on overflow at the same time.
  • The SSD holds a large lookup table so that even models too big for RAM can stream from disk.

Because only 10 of 24,576 experts are needed per token, the odds that a required expert is already on the GPU are high, and the cost of fetching a cold one from RAM is measured in microseconds, not the seconds a naive page-fault implementation would imply. The project publishes a paper and detailed documentation quantifying every part of this pipeline.

Guess, then check

The second half of the performance story is speculative decoding. A small helper model guesses the next few words, and the big model verifies all of them in a single parallel pass, keeping only the correct ones. Because verification is parallel and drafting is cheap, the user sees the same output 1.6-1.8x sooner than conventional autoregressive generation.

Prompt ingestion gets a similar treatment: long inputs are read in chunks of up to 8,192 tokens at a time, which is how Strata sustains over 1,000 tokens per second of prompt processing even on modest hardware. The first message of a long chat takes about a minute per 30,000 tokens; follow-up messages start in seconds.

Pick your size

Strata ships the model in several quantizations, and the installer recommends one based on your RAM. At 32GB, a special “Coder” build — with half the experts removed — fits entirely and retains 91% of the full model’s SWE-bench Verified score, though it is notably weaker outside code, including Chinese and other CJK text. At 64GB, IQ2_XS is the recommended balance; IQ3_S is the best and slowest. Workbench tinkerers with 96GB or more can run Unsloth’s ~4-bit UD-IQ4_XS, and an experimental UD-Q4_K_XL variant streams most of its weights from SSD, writing 7-8.5 tokens per second on a 64GB machine — slow, but proof that even the near-full model fits if you tolerate disk-speed inference.

The hardware requirements are refreshingly ordinary: NVIDIA GeForce RTX 20/30/40/50 series or a range of AMD Radeon cards, 12GB+ VRAM, 32GB+ RAM, ~80GB of disk, Windows or Linux. Community members have pushed it onto Tesla P40s, GTX 10-series cards, Radeon VIIs, and even Intel Arc via a Linux source build.

An OpenAI-compatible endpoint on your desk

What makes Strata more than a benchmark curiosity is the interface. It exposes an OpenAI-compatible API at http://127.0.0.1:8080/v1, plus Anthropic-style (/v1/messages) and Responses-API (/v1/responses) endpoints. Claude Code, Codex CLI, Cursor, and Copilot-style agents can all point at it — meaning a coding agent can drive a 125B-class model with nothing ever leaving your PC.

That last clause is doing a lot of work in 2026. Between the Apple Full Disk Access crackdown on agent overreach, the Gemini Desktop “full access” toggle flap, and state AGs subpoenaing labs over sandbox escapes, the appetite for agentic AI that does not phone home has never been clearer. A model that chats, writes code, reads images, and plugs into your existing agent tooling — entirely on-premise, MIT-licensed, on hardware you already own — is a strong answer to that anxiety. It even ships an MCP server so AI assistants can install, start, and stop the engine themselves.

The fine print

Honesty requires noting the caveats. The headline quantizations are aggressive: Q2_0 is the fastest tier, and two-bit quantization of any 125B model will not match its full-precision self on every task. The “Coder” variant trades general capability for the 32GB RAM footprint. AMD image input on Windows is not supported yet. And a model that loads 35-55GB into RAM will make your PC sluggish or unresponsive for one to three minutes at startup — the README is unusually candid about this.

But the direction is unmistakable. vLLM added Qwen3.8-Flash-Next support within days of release, community reports on r/LocalLLM describe doubled throughput on 3090-class cards, and the Hacker News thread filled with reproducible results. Frontier-class capability is not trickling down to consumer hardware on a yearly cadence anymore — it is arriving on a weekly one. For anyone who wants agents with big-model brains and zero telemetry, Strata just moved the goalposts to the gaming PC already sitting under your desk.