Home/News/Microsoft AI Releases MAI-Cyber-1-Flash for Cyber Defense
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Microsoft AI Releases MAI-Cyber-1-Flash for Cyber Defense

Microsoft AI has released MAI-Cyber-1-Flash, its inaugural model specifically engineered for cyber defense applications. This model is not deployed as a standalone endpoint but operates within MDASH, Microsoft's multi-model agentic scanning harness. MAI-Cyber-1-Flash is a transformer architecture incorporating self-attention mechanisms and sparse Mixture-of-Experts layers. It possesses a total of 137 billion parameters, with 5 billion active parameters, and supports a context length of 256,000 tokens. The model's input and output capabilities are limited to text. It represents a cybersecurity-specialized fine-tune of MAI-Code-1-Flash, a lightweight agentic coding model that is already integrated into GitHub Copilot and VS Code. The release documentation indicates that MAI-Cyber-1-Flash is derived from the MAI-Thinking-1 lineage.

Microsoft evaluated MAI-Cyber-1-Flash using CyberGym, a public benchmark suite comprising 1,507 real-world vulnerability reproduction tasks sourced from 188 OSS-Fuzz projects. The evaluation was conducted at CyberGym's default level 1 configuration, which provides vulnerable source code and a high-level description of the vulnerability. When MDASH, running MAI-Cyber-1-Flash in conjunction with GPT-5.4, was tested, it achieved a score of 95.95% on the CyberGym benchmark. Microsoft has positioned this performance as approximately 12 percentage points higher than Anthropic's Mythos model. A launch chart presented by Microsoft indicates that four competing systems scored between 83.2% and 85.6% on the same benchmark.

When Microsoft initially detailed MDASH in May 2026, the harness had achieved a score of 88.45% on CyberGym, utilizing only generally available models. At that time, this score represented the leading public leaderboard entry, surpassing the next competitor by approximately five percentage points, which scored 83.1%. The research team responsible for the development has stated that this improvement is significant. The CyberGym benchmark is designed to simulate realistic cybersecurity challenges by using actual vulnerabilities found in open-source software projects, making it a robust measure of a model's ability to identify and potentially mitigate security flaws. The inclusion of OSS-Fuzz projects ensures that the tasks are representative of common and impactful security issues faced in the software development lifecycle. The performance metrics suggest a substantial leap in the capabilities of specialized cybersecurity AI models.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next