Interestana
Home/News/BottleCap AI Reduces Thinking Tokens by 37.2% in New Model
MarkTechPost••4 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

BottleCap AI Reduces Thinking Tokens by 37.2% in New Model

BottleCap AI has released ThinkingCap-Qwen3.8-27B, the second model in its ThinkingCap series, which is a fine-tune of the Qwen team’s Qwen3.8-27B model. The primary objective of this new model is to significantly shorten reasoning traces, a critical aspect of large language model (LLM) performance. Across 12 diverse benchmarks, ThinkingCap-Qwen3.8-27B demonstrates an average reduction of 37.2% in the number of "thinking tokens" used. Thinking tokens refer to the internal computational steps an LLM takes to arrive at an answer. This efficiency gain comes with a marginal decrease in macro-average accuracy, which moved from 86.65% in the base Qwen3.8-27B model to 85.79% in ThinkingCap-Qwen3.8-27B, representing a 0.86 percentage point (pp) drop.

The model is designed for deployment and is compatible with vLLM or SGLang frameworks, offering builds in FP8, NVFP4, GGUF, and MLX formats. Access to the model's repository is gated, and while a small-business license is available for commercial use, larger-scale deployment requires a specific agreement with BottleCap AI. The core problem that the ThinkingCap series targets is the tendency for reasoning models to expend more computational tokens than necessary to reach a final answer. BottleCap AI posits that many of these extraneous tokens do not contribute to the accuracy or quality of the output. The first model in the series applied this principle to Qwen3.6-27B, and the current iteration, ThinkingCap-Qwen3.8-27B, maintains a conservative approach. The development team deliberately avoided introducing new knowledge or altering the model's answer style, aiming instead to ensure that core reasoning abilities, instruction following capabilities, and safety behaviors remained largely unchanged from the base Qwen3.8-27B model. The research team placed a particular emphasis on evaluating performance on benchmarks related to mathematics, general reasoning, long-context understanding, and agentic tasks.

All reported benchmark results for ThinkingCap-Qwen3.8-27B were generated using a reasoning effort setting of "xhigh," which is the default configuration for the chat template. This setting ensures a rigorous evaluation of the model's reasoning capabilities. The efficiency gains are evident across all tested benchmarks, with reductions in thinking tokens ranging from a low of 10.7% to a high of 65.5%. Knowledge-intensive and multilingual tasks saw the most substantial reductions. For instance, the MMMLU (Massive Multitask Language Understanding) benchmark experienced a 65.5% decrease in tokens, dropping from 1,656 to 571 tokens. Similarly, MMLU-Pro saw a 57.3% reduction. The GPQA-Diamond benchmark, which tests graduate-level reasoning, saw its token count fall by 43.1%, from 12,772 to 7,267 tokens. The IFBench (Instruction Following Benchmark) utilized 46.4% fewer tokens, with its accuracy remaining nearly constant at 79.71%, a negligible change from the base model's 79.75%.

Performance on long-context retrieval tasks showed improvement. The AA-LCR (Adversarial Adaptive Long Context Retrieval) benchmark saw a 2.25 pp increase in accuracy, rising from 81.75% to 84.00%, while simultaneously reducing thinking tokens by 38.6%. LiveCodeBench v6, a benchmark for code generation, showed a slight increase in accuracy of 0.07 pp, with a 20.3% reduction in thinking tokens. Agentic task performance remained largely stable. The τ²-bench, which evaluates conversational agents, saw a 1.01 pp decrease in accuracy alongside a 30.9% cut in thinking tokens. Terminal-Bench 2.1, another agentic benchmark, lost 0.56 pp in accuracy, a result well within its expected margin of error of ±4.26 pp, while using 10.7% fewer tokens. The most significant accuracy trade-off occurred in the AIME 2026 benchmark, a challenging math competition problem. Here, accuracy dropped by 3.85 pp, from a high of 98.13% to 94.27%, in exchange for a 30.2% reduction in thinking tokens. These results highlight BottleCap AI's success in optimizing reasoning efficiency with minimal impact on core performance metrics across a wide array of tasks.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next