By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Liquid AI Open-Sources Pipette Benchmarking Suite
Liquid AI released Pipette this week, an open-source platform for benchmarking foundation models on edge devices. This initiative was developed in partnership with Artificial Analysis, which served as an independent methodology validator. Pipette's core innovation lies in its approach to evaluating on-device model performance, treating it as a property of the entire deployed system rather than the model in isolation. The platform's unit of measurement is a comprehensive configuration encompassing the model, its quantization method, the runtime environment, and the specific hardware device.
The launch dataset for Pipette includes evaluations across more than 1,000 unique configurations. These configurations combine over 30 different models, various llama.cpp builds for macOS, iOS, Windows, and Android, and context lengths ranging from 256 to 8,192 tokens. The platform measures five distinct on-device performance metrics. Initial verified results have been generated using high-end consumer devices, including a MacBook Pro with an M5 Max chip, an iPhone 17 Pro, and a Galaxy S26 Ultra. A practical demonstration of Pipette's utility shows that two 350M parameter models, when subjected to the same quantization level and tested on the identical phone hardware, retained 78.4% and 33.8% of their decode throughput, respectively, at a context length of 4,096 tokens.
Pipette is designed for broad deployability and accessibility. It is distributed under the Apache 2.0 license and comprises several components: pipette-mgmt for management, pipette-clients for running benchmarks, pipette-scores for data aggregation, a public results dataset, a hosted dashboard for visualization, and native iOS and Android benchmark applications. There are no waitlisting requirements for any of these components. The capability for community-submitted results is currently in a beta testing phase. The platform is intended for any team responsible for shipping models onto hardware they do not own, including solo developers, seed-stage startups, mid-market product teams managing internal device fleets, and large original equipment manufacturers (OEMs), chip vendors, and enterprises that can operate the entire pipeline within their own secure network infrastructure.
The industries that can benefit from Pipette are diverse and include consumer electronics and smartphone OEMs, the automotive sector, industrial automation and robotics, healthcare device manufacturers, financial services, and defense contractors. These sectors frequently face constraints such as latency requirements, privacy concerns, or unreliable connectivity, which necessitate performing inference directly on the device. Pipette's applications extend to critical stages of the development lifecycle, such as model and quantization selection before a development sprint begins, enabling informed decisions based on real-world on-device performance rather than theoretical server-class benchmarks. It also supports hardware selection and optimization, helping teams choose the most suitable devices and configurations for their specific deployment needs.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.