By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Liquid AI Releases Open Multimodal Decision Models
Liquid AI has released two open-weight multimodal models, d1-3B and d1-omni-600M, as part of its d1 decision model family. These models are engineered to provide calibrated, typed answers in a single forward pass, producing zero output tokens, which distinguishes them from traditional generative large language models (LLMs). The d1-3B model is capable of processing both text and images, while the d1-omni-600M model can handle text combined with either images or audio. The primary target for these models is real-time decision-making applications within the NVIDIA technology stack, including DGX servers, RTX workstations, and Jetson edge boards. Both model checkpoints are publicly available on Hugging Face, can be loaded using the Transformers library, and have immediate support for llama.cpp, facilitating broader adoption and deployment. The LFM Open License v1.0 permits free commercial use for entities with annual revenues below $10 million. The d1-omni-600M is an early research release and does not yet have published latency figures.
A decision model, as defined by Liquid AI, differs from a generative LLM. While a generative LLM produces answers token by token, requiring subsequent parsing by external code, a decision model takes a defined state and a set of named questions. It processes this input in a single pass and returns probabilities for each permissible answer. Liquid AI specifies three types of questions these models can address: 'noul' for yes/no inquiries, returning the probability of 'yes'; 'choice' for selecting one option from a named set, providing a full probability distribution; and 'score' for rating a position on an ordered rubric from 2 to 10 levels, using probability weighting. Multiple questions can be processed concurrently within a single state in one model call, with all responses explicitly indicating 'output_tokens: 0'. Liquid AI suggests applications for the d1 models including routing, moderation, intent classification, reranking, LLM-as-a-judge scoring, agent guardrails, and visual inspection tasks. Importantly, neither of the released checkpoints are designed to function as chat models.
The d1-3B model, with 3.12 billion parameters, is built upon LFM2.5-VL-3B, a decoder-only vision-language model. Liquid AI achieved this by averaging the weights of LFM2.5-2.6B with the text backbone of the vision-language model. Subsequently, the company fine-tuned multiple checkpoints using various data mixtures and random seeds, before merging them to create the final model. This development represents a significant step towards more efficient and direct decision-making capabilities in AI systems, moving away from the token-generation paradigm for certain tasks.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.