← All posts / Research

The Other Half of the Memory Wall: AutoArk's Edge0 Streams a 35B MoE From SSD at 20 tok/s

A trained prerouter predicts the next layer's expert routing one token ahead, letting a 35B mixture-of-experts model decode at 20 tok/s inside 2.9 GiB of RAM on a 24GB Mac mini — fully open source.

The Other Half of the Memory Wall: AutoArk's Edge0 Streams a 35B MoE From SSD at 20 tok/s

Running a 35-billion-parameter language model on a consumer desktop used to be a non-starter. Even at aggressive 4-bit quantization the weights alone occupy roughly 19.5 GB, and mixture-of-experts (MoE) sparsity — the trick that makes such models cheap per token — only shrinks the compute needed per token, not the bytes you must hold in memory. On September 16, a team at AutoArk (edge0.ai) posted a paper and an open-source release that attacks exactly this asymmetry. Their streaming inference engine, Edge0, serves a 35B-class MoE at 20 tok/s on a single 24GB Mac mini while keeping peak active memory under 3 GiB.

The paper — “The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction” (arXiv:2609.18063) — is refreshingly direct about why naive offloading fails. Yes, you can leave the expert weights on SSD and page them in on demand. But layer N+1’s experts can only be chosen after layer N’s output exists. The router that decides which experts to fetch is downstream of the computation that needs them, so SSD reads can never start early enough to hide behind the forward pass. The result is the classic offloading death spiral: compute sits idle waiting on storage, and throughput collapses.

The prerouter: predict the routing, then use the prediction

Edge0’s core move is to stop treating routing prediction as a hint and start treating it as the routing itself. A small per-layer head — the prerouter — is trained to predict the next layer’s expert routing one token ahead. During decoding, the prediction is consumed as the actual routing decision, so the set of experts staged from SSD equals the set that gets routed to. Nothing is dropped, nothing is approximated at inference time, and the SSD reads for layer N+1 overlap the forward pass of layer N instead of stalling it.

This is the detail that separates Edge0 from speculative-loading schemes: there is no fallback path and no speculative execution. The trained prediction is the routing. AutoArk reports the prerouter alone buys up to +59% decode throughput, with the gain growing as storage latency, model size, and routed width K increase — precisely the regime where offloading hurts most.

The third mechanism, Recover-LoRA, handles the quality debt. The int4 base is frozen, and LoRA adapters are trained by distillation from the fp16 teacher on the student path — that is, on the exact quantized-plus-replaced-routing pipeline that will actually serve traffic. Adapters stay unmerged, so one read-only base checkpoint can serve multiple adapter sets without re-quantization.

The numbers

On a Mac mini M4 Pro with 24 GB of RAM, the flagship edge0-35b tier — built on Qwen3.6-35B-A3B, 40 layers, 256 experts with 4 active per token — delivers:

MetricValue
Decode speed14.9–17.7 tok/s (paper: 20 tok/s)
Prefill throughput (cold / warm)113 / 140 tok/s
Peak active memory2.9 GiB
Checkpoint on disk (4-bit)~23 GB

The comparison the paper draws against a fully-resident baseline is the striking part: Edge0’s streaming pipeline decodes at 20.4 tok/s inside 2.9 GiB, while a fully-resident int4 configuration holding 18.2 GiB manages only 3.9 tok/s on the same hardware. The apparent paradox — streaming from SSD beating everything-in-RAM — resolves because the resident configuration pushes the machine into swapping, while Edge0’s working set fits comfortably in physical memory.

Quality holds up well under the treatment. Benchmarked with OpenCompass under identical settings, the edge0 pipeline (int4 + adapters + prerouter routing) lands within 3.9 points on average of its fp16 teacher: AIME 2026 at 86.6 vs 92.7, HumanEval at 90.9 vs 95.1, GPQA-Diamond at 79.8 vs 81.8, MMLU-Pro at 81.0 vs 84.6, IFBench at 57.9 vs 61.7. A smaller edge0-8b tier (built on the Ling 3.0 bailing hybrid, 128 experts, prerouter K=8) fares even better at 2.8 points average — and its MMLU-Pro score of 70.1 actually beats the fp16 base’s 65.8.

Why it matters

The memory wall is the binding constraint on edge AI. Datacenter-scale inference hides it with HBM and expert parallelism; consumer devices cannot. Edge0’s contribution is showing that the wall’s “other half” — the routing dependency that blocks prefetch — yields to a learned predictor, not just to faster storage or bigger caches.

The implications run in several directions:

  • Phone-class deployment becomes plausible. A 35B MoE running in under 3 GiB of active memory fits hardware that could never hold 20+ GB of weights. The 8B tier runs in ~1.0 GiB.
  • Commodity batch serving. Because the base stays read-only and adapters are separate safetensors, one machine can serve many LoRA-customized variants from a single checkpoint on disk.
  • The recipe generalizes. Edge0 is framed as a framework, not a single model hack: SSD expert offload, prerouter, and Recover-LoRA are packaged as extensible abstractions with an isolated MLX backend (Apple Silicon today, CUDA on the roadmap).

There are honest caveats. This is a preview release; agentic capability — tool use, multi-step planning, long-horizon autonomy — is explicitly weak in this build, with the full release promising substantial strengthening. The MLX backend currently targets Apple Silicon only, and long contexts grow the KV cache beyond the headline 3 GiB figure. Interactive speed at 15–20 tok/s is usable but not luxurious for a 35B model, and the benchmark methodology (3.3k-token prefill, 200 timed decode tokens, two runs) leaves room for independent replication.

Still, the shape of the result is what matters. In a week where the industry conversation is dominated by trillion-dollar valuations and hundred-thousand-chip clusters, Edge0 is a reminder that the frontier of efficiency research is moving fast too — and that a 35B-parameter model streaming its experts from a consumer SSD, with routing predicted by a trained head one token ahead, is now an Apache-2.0 download away.