Interestana
Home/News/Underdog Saluki 27B Beats Original Qwen3.8-27B at Tool Calling
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Underdog Saluki 27B Beats Original Qwen3.8-27B at Tool Calling

Underdog, a division of Conway Research, has released Saluki 27B, a highly compressed 2-bit GGUF model derived from Qwen3.8-27B, under the permissive Apache 2.0 license. This new model is specifically engineered for on-device AI assistants and achieves a remarkable file size of 7.89 GB, a substantial reduction from the 54 GB required by the full BF16 version of Qwen3.8-27B. The primary objective of this compression was to preserve and enhance tool calling capabilities, a critical function that transforms a standard chat model into a more versatile AI agent. Developers can now leverage a 27B-parameter agent model that operates seamlessly within the stock llama.cpp framework, enabling broader accessibility and deployment on consumer hardware.

Saluki 27B is built upon a multi-layered approach to quantization and fine-tuning. The foundation is Qwen3.8-27B, a dense 27 billion parameter model developed by the Qwen team, featuring 64 layers and incorporating Gated DeltaNet linear attention alongside gated attention, with native support for up to 262,144 tokens. The second layer of development involved integrating ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF, which employs Group-wise Scalar Quantization (GSQ) to learn precise low-bit scalar grids for each tensor and Residual-based Channel Optimization (RCO) to assign quantization types within a fixed size budget. ISTA's smallest variant, IQ2_XS, typically occupies 8.4 GB at 2.50 bits per weight. Underdog's proprietary third layer further refined this process, shrinking the file to 7.89 GB and specifically optimizing for tool calling performance. This final pass resulted in a file named IQ2-mix, which includes an imatrix tag, though Underdog has not disclosed the complete methodology behind this optimization.

Performance evaluations indicate that Saluki 27B maintains a strong overall capability, achieving 96% of the performance of the full Qwen3.8-27B model across nine benchmarks. Its most significant advantage lies in parallel tool calls, where it surpasses the original model by a considerable margin, handling 42 concurrent calls compared to the full model's 35, representing a 120% retention rate in this specific function. However, the aggressive quantization does lead to a noticeable decrease in performance on certain complex reasoning tasks. For instance, in the AIME 2025 benchmark, which focuses on competition mathematics, Saluki 27B achieved a score of 79.2%, compared to 96.7% for the full model, indicating approximately an 82% retention rate. This suggests a trade-off where enhanced agentic capabilities are prioritized over performance in advanced mathematical and multi-step reasoning problems, with a drop of 12 to 18 percentage points in these areas.

For developers and users seeking efficient, on-device AI agents, Saluki 27B presents a compelling option. Its ability to run on stock llama.cpp, with full GPU offload capabilities, makes it highly accessible. Furthermore, an optional vision add-on, available in 629 MB or 928 MB sizes, further expands its utility by enabling visual understanding. The model's success in outperforming the larger, uncompressed original model in critical agent functions like tool calling, while fitting into a fraction of the storage space, highlights the advancements in efficient AI model deployment. The trade-off in specific reasoning benchmarks is a key consideration for applications requiring high precision in areas like competitive mathematics or complex problem-solving.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next