Interestana
Home/News/Liquid AI Releases LFM2.5-VL-3B On-Device Vision-Language Model
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Liquid AI Releases LFM2.5-VL-3B On-Device Vision-Language Model

Liquid AI released LFM2.5-VL-3B yesterday, a 3.1 billion-parameter vision-language model engineered for on-device deployment. This model is capable of reading digital screens across mobile, web, and desktop interfaces. It can also ground objects to specific coordinates, parse complex documents and charts, and invoke tools based on either text or image input. Liquid AI reported an average score of 69.4 across 28 vision benchmarks, a performance that matches the 4.7 billion-parameter InternVL-3.5-4B model and trails the Qwen3.5-4B model, also 4.7 billion parameters, by 0.7 points. The LFM2.5-VL-3B is designed for low latency, operating without complex reasoning to provide direct answers. It requires approximately 3 GB of memory and can decode 228 tokens per second on an Apple M5 Max chip. The model is deployable in four formats: native, GGUF, ONNX, and MLX, with day-one runtimes supported by llama.cpp, MLX, vLLM, SGLang, and ONNX. The LFM Open License v1.0, which governs its use, is based on Apache-2.0 with a modification: free commercial use is permitted until a company's annual revenue reaches $10 million USD. This allows independent developers, startups, and small to medium-sized businesses below this revenue threshold to utilize the model commercially without cost. Enterprises exceeding this revenue limit must negotiate a separate commercial license with Liquid AI. Research, educational, and non-profit organizations can use the model without any revenue restrictions. The model is targeted for industries including consumer electronics, automotive, industrial and robotics, financial services, healthcare, and e-commerce, as well as for Quality Assurance (QA) and Robotic Process Automation (RPA) vendors focused on automating graphical user interfaces (GUIs). Potential applications include on-device screen agents, GUI test automation, converting PDFs to structured text with layout information, processing invoices and receipts for OCR, near-real-time object detection in vehicles, offline translation of menus and road signs, and multi-image comparison tasks. LFM2.5-VL-3B represents an advancement over its predecessor, LFM2-VL-3B, enhancing capabilities in four key areas. Its screen and UI understanding performance is notable, achieving an average of 80.7 on the ScreenSpot-v2 benchmark, with scores of 78.7 on desktop, 81.2 on mobile, and 82.2 on web interfaces. Liquid AI positions this performance against Gemma-4-E4B at 51.2 and Qwen3.5-4B at 78.5, with InternVL-3.5-4B leading at 84.1. Function calling is a new addition to the vision-language line, with performance on the ToolSandbox benchmark improving from 26.4 to 59.5, and BFCL v4 moving from an unspecified previous score to 2.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next