← All posts / Tools

3.5x to 10x Faster Image and Video Inference: Nunchux Exits Stealth With the Modelverse

Born from MIT HAN Lab and CMU research, Nunchux launches the Modelverse with up to 10x faster multimodal inference — FLUX.2-klein in 0.08s, MiniMax-H3 video in 10.4s — without losing quality.

3.5x to 10x Faster Image and Video Inference: Nunchux Exits Stealth With the Modelverse

For most of the last decade, AI inference infrastructure has been optimized around one shape of workload: autoregressive language decoding, one token at a time. This week, a startup spun out of two of the most-cited academic labs in efficient generative modeling came out of stealth arguing that this foundation is the wrong one for the next wave of AI workloads — and shipped a product to prove it.

Nunchux AI, built by a team drawn from MIT’s HAN Lab and Carnegie Mellon’s Generative Intelligence Lab, officially launched its Modelverse this week: a serving platform for image, video, and world models that the company claims delivers the fastest, cheapest, and highest-fidelity multimodal inference available today. The headline numbers are hard to ignore: 3.5x faster on FLUX.1-schnell, 10x faster on FLUX.2-klein-4B, 4.9x faster on MiniMax-H3 video generation, and 4.4x faster on LTX-2.5-fast.

The physics of the problem

The core insight behind Nunchux is deceptively simple. Language models and visual models have fundamentally different computational profiles. In autoregressive text decoding, each step generates a single token, and the bottleneck is usually memory bandwidth — the model’s weights must be streamed through the accelerator for every token produced. Serving stacks built for LLMs (vLLM, SGLang, TensorRT-LLM) are all tuned around this memory-bound regime.

Image and video diffusion models invert that profile. At each generation step, an entire spatial or spatiotemporal latent tensor is processed in parallel — dense, arithmetic-heavy compute. The bottleneck migrates from memory to raw compute, and the scheduling tricks, batching strategies, and economics that work for language simply do not transfer. Video is the extreme case: the company notes it takes roughly two orders of magnitude more compute than a single image, “but nobody wants to wait two orders of magnitude longer.”

Rather than tuning an existing stack, Nunchux built its own from the ground up: the Nunchux Model Optimizer and the Nunchux Engine, designed as a single system so that gains at one layer create headroom for the next. The stack runs across NVIDIA GPUs, AMD GPUs, and even mobile devices.

The research lineage

What makes Nunchux more credible than the average inference startup is its research pedigree. The underlying techniques — SVDQuant, Sparse VideoGen, DeltaQuant, and Radial Attention — came out of a decade of collaboration between MIT HAN Lab (best known for the TinyML and efficient-deep-learning lineage) and CMU’s Generative Intelligence Lab, work that has accumulated more than 200,000 citations.

SVDQuant alone has been downloaded 4 million times and adopted as the standard quantization algorithm inside NVIDIA TensorRT, AMD Quark, vLLM, SGLang, Diffusers, and ComfyUI — meaning that if you have run any optimized open image model in the past year, there is a decent chance Nunchux’s research was already somewhere in your serving path.

The company argues speed cannot come from a single trick. Its SVDQuant pipeline runs FLUX.1-schnell 3x faster while improving the ImageReward quality metric from 0.938 to 0.966 — faster outputs that are also measurably better-looking, not degraded approximations.

What shipped: the Modelverse

The Modelverse catalog is organized into two performance tiers: Radical Speed for latency-critical applications and Radical Value for cost-critical ones. Alongside Nunchux-optimized open models, the platform offers Partner APIs from foundation labs, including Google Veo 3.1, HeyGen Avatar 5, and Kling V3, with more than 30 optimized and partner APIs in total.

The measured benchmarks the company published alongside launch:

ModelTaskSpeedupNunchuxBaseline
FLUX.1-schnellimage3.5x0.143 s0.5 s
FLUX.2-klein-4Bimage10.0x0.08 s0.8 s
Qwen-Imageimage2.6x2.24 s5.9 s
Ideogram v4image2.4x6.94 s17 s
LTX-2.5-fastvideo4.4x6.6 s28.9 s
MiniMax-H3video4.9x10.4 s51.3 s

A sub-100-millisecond FLUX.2 generation is a meaningful threshold: it moves image generation from “submit and wait” into the territory of genuinely interactive creative loops, where each output can inform the next prompt in real time.

Quality is enforced with a gate, not a promise. Every model is checked against its original output before entering the catalog, and any optimization that visibly degrades fidelity keeps the model out.

Why the economics matter

Nunchux’s framing of the cost-quality-latency triangle is worth pausing on. Multimodal generation is inherently iterative: users generate, look, edit, and try again, frequently without knowing what they want until they see it. Cheap attempts widen the design space; expensive ones force premature commitment. Latency decides whether the result arrives “while the idea is still yours.”

If these speedups hold in production, the practical consequence is that the per-iteration cost of visual AI development collapses — which in turn changes what products become viable. Real-time image editing agents, video previews embedded in search results, and world-model simulation loops all become cheaper by roughly an order of magnitude.

The company is backed by lead investors Emergence Capital and E14 Fund, with participation from Archerman Capital, Scale Asia Ventures, Tectonic Ventures, KungHo Fund, and Plaid Matrix Fund, plus angels from Black Forest Labs, Reve, HeyGen, Pika Labs, Together AI, and founding members of Meta Superintelligence Labs. Demand is reportedly ahead of capacity, so onboarding runs through a waitlist, with $10 in free credits for new accounts.

The bigger picture

Inference optimization has quietly become one of the most strategically contested layers of the AI stack. NVIDIA keeps absorbing optimization research into TensorRT; Decart was reportedly valued at $6 billion in an Anthropic acquisition discussion earlier this month; and every hyperscaler is racing to make multimodal generation cheap enough for consumer-scale products. Nunchux represents the academic counter-punch: the labs that wrote the papers the industry now depends on, commercializing their own stack directly.

The open question is defensibility. Compiler tricks and quantization schemes get commoditized quickly — as SVDQuant’s own integration into TensorRT and vLLM demonstrates. The bet Nunchux is making is that staying at the frontier of the quality-cost-latency Pareto surface, across both NVIDIA and AMD silicon and down to mobile, is a durable position rather than a temporary one. If the 10x speedups hold and the quality gate stays honest, the Modelverse just became the reference point for what multimodal inference should cost.