Interestana
Home/News/Liquid AI Ships On-Device Agentic Model LFM2.5-2.6B
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Liquid AI Ships On-Device Agentic Model LFM2.5-2.6B

Liquid AI released LFM2.5-2.6B, an agentic large language model designed to run entirely on-device, enabling local planning, tool calling, and multi-step task execution across devices such as phones, laptops, PCs, and robots. The model features 2.69 billion total parameters and boasts a substantial context window of 131,072 tokens, with a vocabulary size of 128,000 tokens. Its pre-training phase utilized approximately 34 trillion tokens. Liquid AI has provided two distinct checkpoints: LFM2.5-2.6B-Base, intended for fine-tuning by developers, and LFM2.5-2.6B, which has undergone post-training specifically for agentic workloads. A key advantage of this model is its on-device inference capability, which ensures that data remains local, thereby enhancing privacy and security while reducing the marginal cost of each operational run to near zero. Liquid AI reports that the model's tool-use and instruction-following scores are competitive with models that are nearly four times its size. The model's deployability is confirmed, with both checkpoints publicly available on Hugging Face under the lfm1.0 license. The model weights are distributed in native, GGUF, MLX, and ONNX formats, with immediate support for popular inference engines including llama.cpp, vLLM, SGLang, and LM Studio. This release empowers solo developers and startups to pilot the model on existing hardware, with decoding speeds of 220 tokens per second observed on an M5 Max chip using under 2.5 GB of RAM. For mid-market teams, self-hosting on a single GPU is feasible, where one NVIDIA H100 SXM5 GPU can serve approximately 1.3 billion tokens daily. Enterprises and Original Equipment Manufacturers (OEMs) can deploy these weights across device fleets using GGUF and ONNX formats. Fine-tuning can be efficiently performed using LoRA techniques with libraries such as TRL and Unsloth. Liquid AI is targeting a broad range of industries, including automotive, consumer electronics, industrial robotics, healthcare, financial services, e-commerce, and defense. The model is particularly beneficial for regulated environments and air-gapped systems, as it eliminates the need to send prompts to third-party APIs. Recommended applications for LFM2.5-2.6B include agentic workloads, tool utilization, data extraction, Retrieval-Augmented Generation (RAG), and long-context workflows. Specific practical applications include developing on-device assistants, performing offline document triage on inputs exceeding 128,000 tokens, extracting information from forms and invoices, parsing commands for robotics, and implementing background agents that operate continuously without incurring per-token costs. Liquid AI explicitly advises against using the model for agentic coding tasks or knowledge-intensive applications where its performance might be suboptimal compared to larger, cloud-based models.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next