Interestana
Home/News/Reflection AI Releases Beam: 501B Open-Weight MoE Model
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Reflection AI Releases Beam: 501B Open-Weight MoE Model

Reflection AI has introduced Beam, its inaugural open-weight model, on March 18, 2024. Beam is a sparse Mixture-of-Experts (MoE) architecture boasting a total of 501 billion parameters, with 23 billion active parameters per token. This model is specifically engineered for coding, reasoning, and agentic workloads. According to the Reflection AI team, Beam is positioned to directly compete with larger open models such as GLM 5.2, while demonstrating a significant advantage in inference compute efficiency, using 3 to 4 times less compute on reasoning benchmarks. Currently, Beam is not available for self-hosting but is undergoing final red-teaming. Early access is being managed through a waitlist on the Reflection platform. Beam functions as a general agent model, developed from the ground up by Reflection AI, with a primary focus on enterprise-level coding and agentic applications. Reflection AI frames Beam as a significant advancement in the Western open-weight model landscape. The research team acknowledges that while Kimi K3 may still lead in raw capability, Beam's key selling proposition is its efficiency during inference. Users can adjust a "reasoning effort" parameter, where lower settings yield concise answers and higher settings enable more extensive reasoning for complex tasks. This allows teams to tailor the model's effort to the specific difficulty of a task and their available compute budget. The pretraining of Beam involved an extensive dataset of 23.8 trillion tokens, sourced from the web, public repositories, and proprietary licensed datasets. The Reflection AI team reported that their curation process removed approximately 95% of raw internet tokens, while retaining about 1.8 trillion high-quality tokens that might have been discarded by conventional filtering methods. The model's architecture integrates local and global attention mechanisms with finely routed experts. Its load balancing strategy builds upon DeepSeek-V3's auxiliary-loss-free balancing, incorporating a cosine decay for expert-bias updates. This approach ensured that the most utilized expert experienced a load of only 1.04 times the average by the conclusion of pretraining. Throughout the 52 layers, residual norms were maintained within bounds through depth-based scaling, SandwichNorm, attention gating, and FP32 residual accumulation. The entire pretraining process was completed in under four weeks, utilizing 6,144 NVIDIA GB300 NVL72 GPUs. The reported goodput reached 92.3% near the end of the training period, with nine semi-automatic rewinds implemented. During mid-training, the effective context window was extended to 1 million tokens.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next