Liquid AI Ships LFM2.5-2.6B: Open-Weight Agentic Model That Runs on Your Phone
Liquid AI's 2.6B parameter model plans, calls tools, and runs multi-step agent workflows entirely on-device at 220 tokens/s — no cloud, no GPU required.
Liquid AI released LFM2.5-2.6B on August 4, 2026, and it represents something genuinely different in the current AI landscape: a small, open-weight model purpose-built for on-device agentic workflows. While the industry’s attention has been consumed by frontier models demanding hundreds of billions of dollars in compute infrastructure, Liquid AI has been quietly building models designed to run on the phone in your pocket — and the results are turning heads.
The pitch is straightforward. LFM2.5-2.6B is a 2.69-billion-parameter model that can plan, call tools, and execute multi-step agent tasks entirely on local hardware. It needs under 2.5 GB of memory. It decodes at 220 tokens per second on an Apple M5 Max, 113 tokens per second on a Ryzen AI Max+ 395, and still holds 30 tokens per second on a phone. No cloud API calls. No per-token costs. No data leaving the device.
Why On-Device Agents Matter
The economics of cloud-based AI agents have always been a constraint. Every token generated by a frontier model behind an API costs money, which means developers must carefully budget how much reasoning, exploration, and iteration their agents are allowed to do. This fundamentally limits what agent architectures can attempt. Background tasks that burn through millions of tokens — continuous monitoring, speculative exploration, massive parallelization — become financially unsustainable.
Liquid AI’s argument is that removing the per-token cost changes the equation entirely. When inference is free because it runs on local hardware, agents can be deployed everywhere, running continuously, without worrying about the bill. A fleet of agents on a rack of commodity machines, or even on individual phones, can work around the clock at zero marginal cost.
The privacy angle matters too. On-device agents process data locally, which eliminates the need to send sensitive information to a third-party API. For healthcare, finance, legal, and enterprise use cases where data residency is non-negotiable, this is not a convenience — it is a requirement.
Architecture and Training
LFM2.5-2.6B is pre-trained on approximately 34 trillion tokens. The vocabulary was doubled to 128K — extending the existing tokenizer in place rather than retraining from scratch — to improve support for non-Latin scripts. Mid-training includes a dedicated 128K context-extension phase so the model can handle the long inputs that agentic workloads demand.
The post-training pipeline is where things get interesting. It consists of four stages:
1. Supervised Fine-Tuning (SFT): Two consecutive SFT stages — broad coverage first, then targeted shaping on priority skills like agentic tasks, reasoning, and tool use. The SFT training mix is roughly seven times the size of the one used for Liquid AI’s larger LFM2.5-8B-A1B model, with heavy weighting toward tool use, web search, software engineering, and agent traces.
2. Teacher Specialization: From the shared SFT checkpoint, Liquid AI trains one expert per target domain through focused SFT followed by reinforcement learning with verifiable rewards (RLVR). The specialists cover instruction following, math, knowledge with hallucination control, code, tool use, and long context. Training them separately allows each expert to optimize deeply without competing updates from unrelated objectives.
3. Multi-Domain On-Policy Distillation (MOPD): The specialized experts serve as teachers, distilling their capabilities into a single student model. Unlike off-policy distillation, where a student learns from trajectories generated by another model, MOPD lets the student roll out under its own policy. Each prompt is routed to the domain-appropriate teacher, which provides token-level feedback. Because the teachers branch from the same SFT checkpoint as the student, the feedback stays close to the student’s distribution, preventing training instability.
4. Agentic RL: The final stage trains the model inside real agent environments. Multi-turn agentic reinforcement learning runs through actual agent harnesses where the model tackles realistic productivity tasks — research, writing, coding, data analysis, document management, tool use, and workflow automation. The model is optimized with GRPO using outcome-based rewards that combine an LLM-as-a-judge rubric, programmatic checks, and a hard safety gate. Critically, training directly inside harnesses like Hermes Agent and OpenClaw exposes the model to their tools, system prompts, and interaction patterns.
Benchmark Performance
Despite being the smallest model in Liquid AI’s comparison set, LFM2.5-2.6B is remarkably competitive. It leads on every instruction-following benchmark tested (IFBench, Multi-IF, IFStruct) and nearly every tool-use benchmark, trailing only Qwen3.5-9B on BFCLv4. On agentic tasks, it outperforms both Gemma models across the board and trades closely with the larger Qwen models.
Specific highlights: 51.87 on AIME25 (beating both Gemma models and Qwen3.5-4B), 56.88 on BFCLv4 (function calling), 77.83 on ToolSandbox, and 68.22 on PinchBench. The one area where larger models maintain a clear edge is coding — LiveCodeBenchv6 scores show the 8B and 9B models pulling ahead.
For a 2.6B model competing against 5B, 8B, and 9B alternatives, these results validate Liquid AI’s efficiency-first architecture. The model is not trying to be the best at everything — it is trying to be the best agent that fits on a phone, and on that metric it delivers.
Inference and Deployment
LFM2.5-2.6B ships with day-one support across the full inference ecosystem:
- llama.cpp — GGUF checkpoints for efficient edge inference
- MLX — Optimized for Apple Silicon
- vLLM — GPU-accelerated serving for production throughput
- SGLang — GPU-accelerated serving for production
- ONNX — Cross-platform inference across diverse accelerators
On GPU, the model reaches nearly 15,000 output tokens per second at high concurrency on a single H100, translating to roughly 1.3 billion tokens per day. That throughput makes it viable not just for edge deployment but for high-volume production serving where cost per token matters.
Setting up a local agent takes two steps: serve LFM2.5-2.6B behind an OpenAI-compatible endpoint, then point an agent harness at it. It works out of the box with Hermes Agent, OpenClaw, and Pi — no modifications needed.
The Bigger Picture
LFM2.5-2.6B arrives at a moment when the AI industry is bifurcating. On one side, companies like OpenAI, Anthropic, and Google are pouring enormous resources into ever-larger frontier models running on massive cloud infrastructure. On the other, a growing movement is building toward efficient, deployable models that bring intelligence to the edge.
Liquid AI, founded by Ramin Hasani, Mathias Lechner, and their team from MIT CSAIL, has been pursuing this efficiency-first vision since 2023. The LFM architecture — based on liquid neural networks and designed for compute efficiency — is their bet that the future of AI is not only about scale but about where intelligence lives.
The model is fully open-weight, downloadable from Hugging Face with no restrictions. Developers can fine-tune, modify, and deploy it freely. For the open-source community, this represents another significant addition to the growing catalog of capable small models — and one that is specifically optimized for the agentic use cases that define the current frontier of AI application development.
Whether LFM2.5-2.6B becomes the default choice for on-device agents remains to be seen. But the combination of agentic capability, open weights, extreme efficiency, and real-world benchmark performance makes it a compelling option for anyone building agents that need to run locally, privately, and at scale.
Sources
- [1] https://www.liquid.ai/blog/lfm2-5-2-6b
- [2] https://venturebeat.com/technology/no-cloud-no-gpus-no-problem-liquid-ais-new-model-lfm2-5-2-6b-brings-powerful-ai-agents-to-devices-as-small-as-a-raspberry-pi
- [3] https://www.marktechpost.com/2026/08/06/liquid-ai-lfm2-5-2-6b-on-device-agentic-model/
- [4] https://huggingface.co/blog/LiquidAI/lfm2-5-2-6b