← All posts / Models

Eight Milliseconds, Zero Tokens: Liquid AI Open-Sources Its d1 Decision Models for the Edge

Liquid AI releases open-weight d1-3B and multimodal d1-omni-600M — models that answer in one forward pass with no output tokens, running on everything from an RTX 4090 to a Jetson Orin Nano.

Eight Milliseconds, Zero Tokens: Liquid AI Open-Sources Its d1 Decision Models for the Edge

For two years the AI industry’s default answer to every problem has been the same: spin up a large language model, prompt it, and parse the tokens it writes back. Liquid AI is betting that a growing class of workloads never needed generated text in the first place. On October 7, 2026, the Cambridge, Massachusetts-based company released Open d1, two open-weight models from its d1 decision model family: d1-3B and the experimental multimodal d1-omni-600M. Both are available now on Hugging Face, and both share an unusual property — they produce an answer in a single forward pass without generating a single output token.

What a decision model actually is

Decision models are a new category of AI system purpose-built for structured choices: classification, routing, scoring, moderation, intent detection. Instead of writing a paragraph of prose that another program must then parse, a decision model maps an input directly onto a fixed set of outcomes and returns calibrated probabilities across them. One API call in, one structured answer out.

The efficiency argument is straightforward arithmetic. A generative model answering a routing question must write out its reasoning token by token, paying for every one of them and waiting for the full sequence to complete. A decision model collapses that entire loop into one pass through the network. Liquid AI reports that d1-3B answers a question in 8 milliseconds on an NVIDIA RTX 4090, and 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50 ms even on the tiny Jetson Orin Nano — fast enough for real-time decisions on the smallest edge hardware NVIDIA ships. On an AMD MI325X, the same model handles a packed batch of 64 states at 1,106 per second.

This places Open d1 inside a rapidly crowding category. Cloudflare shipped its open-weight Clef decision models on October 1; Amazon released Strands Decider 2B the same day; OpenAI’s Decisions API entered public beta on October 7. But Liquid AI’s release is differentiated on two axes: raw speed at the edge, and multimodality.

Two models, two backbones

The two checkpoints come from very different lineages. d1-3B is trained from LFM2.5-VL-3B, Liquid AI’s latest decoder-only vision-language model, and accepts text and images as input. The company says it scores 48.57 on the Decision Index v0.2.1 (public split) — ahead of every model under 10B parameters, and on par with Decider 35B-A3B, a decision model twelve times its size.

The training recipe is notably unglamorous. Liquid AI averaged the weights of LFM2.5-2.6B and the text backbone of LFM2.5-VL-3B to build the base, fine-tuned multiple checkpoints with different random seeds and data mixtures, then merged them again. The team notes that training on long inputs, shuffling answer options, and fixing shortcuts in the data “made a bigger difference than more advanced techniques” — a candid admission that data hygiene, not architectural novelty, drove most of the gain.

d1-omni-600M is the more experimental of the two. Built from LFM2.5-Encoder-350M, a bidirectional encoder, it adds a FastConformer audio encoder (trained with an adapter against a frozen text backbone) and a vision encoder drawn from LFM2.5-VL-450M (attached via adapter plus LoRA updates active only when images are present). The result handles text plus image, or text plus audio — three modalities in a 600-million-parameter package. It scores 15.95 on the same Decision Index, and Liquid AI is explicit that it is an early research checkpoint under active development.

The benchmark picture

On seven public text benchmarks spanning reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding, d1-3B posts a mean of 82.9 — the highest in Liquid AI’s comparison table, ahead of Decider 4B at 81.1. The wins are not uniform: d1-3B leads decisively on SQuAD 2.0 (85.3 vs 76.0) and PAWS-X paraphrase detection (76.9 vs 69.8), while trailing Decider 4B on BoolQ, XNLI, and MASSIVE intent classification.

d1-omni-600M is arguably the more interesting story despite the lower headline numbers. At a quarter of Decider 2B’s parameter count it reaches a mean of 78.4 versus 77.1, and it posts the single highest score anywhere in the table on toxicity detection (Civil Comments: 95.8) and paraphrase identification (PAWS-X: 79.5). For footprint-constrained deployments — a smart camera, a wearable, an industrial sensor — trading a few points of intent-classification accuracy for a 600M multimodal model is a rational deal.

The vision story is thinner. Liquid AI validated that d1-3B retains its VLM backbone’s vision capabilities and that d1-omni-600M functions across all three modalities, but the Decision Index’s private vision split is not reported in this release. Audio decision benchmarks, the company acknowledges, are “currently an open problem” — an honest gap in a field this young.

Why the edge matters here

The strategic weight of the release sits in the deployment matrix. Open d1 runs the full NVIDIA stack — DGX in the data center, RTX workstations, Jetson modules at the edge — with day-one llama.cpp support across Apple silicon, AMD, Qualcomm, and NVIDIA, including NVFP4 quantization. On an Apple M5 Pro, d1-3B answers one question in 30 ms and chews through 78 packed states per second.

To demonstrate what real-time decisions look like in practice, Liquid AI built ten demos running d1-3B in a loop over live camera input — gesture-controlled games, live content moderation — each reading an answer from a single pass per frame, all runnable in a Hugging Face Space with no setup. In collaboration with NVIDIA, the company also shows d1-3B navigating an environment in Isaac Sim with the model served on a Jetson in a hardware-in-the-loop configuration.

That last demo points at the near-term market: robotics and embodied agents constantly make discrete choices (turn left, grip, stop, flag this frame) that today are either hand-coded heuristics or grotesquely over-provisioned LLM calls. A 3B model answering in 8–50 ms on-device fits that slot exactly. Reddit commenters were quick to spot another application — NPC decision-making in video games, where a finite action list and millisecond latency budgets are the native conditions.

The bigger picture

Three trends converge in this release. First, the specialization wave: October 2026 opened with eight models in three days and not one frontier generalist among them — the action has moved to purpose-built systems. Second, cost per completed task is displacing cost per token as the routing metric teams actually optimize, and a model that emits zero output tokens is the limit case. Third, open weights as competitive strategy: Liquid AI’s proprietary d1 API launched a week earlier, and the company is now giving away the edge-sized variants to seed adoption — the same land-grab logic driving Cloudflare, Amazon, and Reflection AI.

Open d1 is a research checkpoint, not a finished product category. The omni model is experimental, audio benchmarks don’t yet exist, and a 48.57 Decision Index score says little about how these models behave under adversarial distribution shift. But as a statement of direction it is unusually clear: the future of AI at the edge is not a smaller chatbot. It is a new kind of model that never chats at all — it just decides.