By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Aleph Alpha Releases Kolibri: 78.1B Parameter Open-Weight Model
Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model designed for both English and German languages. This model boasts a total of 78.1 billion parameters, but crucially, it activates only 3.46 billion parameters, representing 4.4% of the total, per token. This efficient activation strategy allows for reduced computational requirements during inference. Kolibri supports an extensive context window of up to 1,048,576 tokens, significantly larger than many contemporary models. Users have the flexibility to adjust the "reasoning effort" for each specific request, enabling fine-grained control over performance and resource utilization. The model is distributed under the permissive Apache 2.0 license and is available on Hugging Face, making it accessible for a wide range of applications. Aleph Alpha specifically targets sovereign deployment within regulated sectors, including public administration, industrial applications, and aerospace, emphasizing data privacy and control.
The technical specifications of Kolibri indicate its deployability on modern hardware. A full FP8 checkpoint for the model occupies approximately 78GB of storage. It can be run on a single NVIDIA B200, B300, or H200 GPU, or alternatively on two H100 SXM5 GPUs. The model is served using vLLM, an optimized inference engine, and includes dedicated parsers for Kolibri's reasoning capabilities and tool-calling functionalities. Kolibri-1, as it is also known, is a bilingual transformer developed entirely by Aleph Alpha's teams in Germany. The company asserts full control over the entire development pipeline, encompassing data curation, architecture design, training infrastructure, post-training refinement, and evaluation processes. The training operations were conducted across infrastructure located in Germany and Finland. The model's design and development adhere to stringent European regulatory frameworks, including the EU General-Purpose AI Code of Practice, the EU AI Act, and the General Data Protection Regulation (GDPR). Aleph Alpha, as a signatory to the AI Code of Practice, implements a data pipeline that proactively redacts personal data before the training phase commences, ensuring compliance with privacy mandates.
Kolibri's architecture features a sparse expert design combined with a hybrid attention mechanism. The model is structured with 50 transformer blocks, each having a model width of 2,560. Every MoE layer employs a sigmoid router to score all 384 available experts. Each token is then routed to the top 6 experts, with one shared expert always being utilized. Expert load balancing is managed through an Exact Quantile Balancing algorithm and a Load-Error Injection technique to ensure efficient distribution of computational tasks. The attention mechanism utilizes grouped-query attention, featuring 48 query heads and 4 key-value (KV) heads. Notably, every fifth transformer block implements full attention without positional encoding. The remaining 40 blocks utilize a sliding-window attention mechanism over the preceding 512 tokens, enhanced with Rotary Positional Embeddings (RoPE). The sliding-window layers maintain a fixed-size KV cache, meaning only 10 layers dynamically grow with the context length. Aleph Alpha reports that this hybrid attention approach supports sequences that are up to four times longer than those manageable by a full-attention model at equivalent computational loads. Furthermore, a specialized tokenizer has been developed specifically for the German language, optimizing its performance for German text processing.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.