By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Sakana AI Releases Fugu-Cyber Cybersecurity Model
Sakana AI released Fugu-Cyber (model ID fugu-cyber-v1.0) this week, an addition to its Fugu orchestration family specifically tuned for cybersecurity reasoning. This model represents a third endpoint on the Fugu orchestrator, which Sakana launched approximately one month prior. Sakana reports that Fugu-Cyber achieved a success rate of 86.9% on the CyberGym benchmark and 72.1% on the CTI-REALM benchmark. The company states these results are comparable to other cyber-focused frontier models, including OpenAI's GPT-5.5-Cyber and Anthropic's Claude Mythos Preview.
The CyberGym benchmark, developed by UC Berkeley, evaluates agents on 1,507 real-world vulnerabilities across 188 OSS-Fuzz projects. Its primary task requires an agent to receive a vulnerability description and an unpatched codebase, then generate a proof-of-concept that successfully crashes the pre-patch build but not the post-patch build, with verification steps designed to prevent manipulation. CTI-REALM, a detection-engineering benchmark from Microsoft, uses 37 public threat reports. Agents must map MITRE ATT&CK techniques, analyze telemetry, refine KQL queries, and produce validated Sigma rules, with scoring covering Linux endpoints, Azure Kubernetes Service, and Azure cloud environments. Together, these benchmarks span the process from identifying and proving bugs to translating threat intelligence into actionable detections.
Sakana's reported 86.9% on CyberGym represents a marginal improvement over existing frontier models. When the CyberGym researchers initially published results, the top agent-model pairing achieved around 20%. More recently, Anthropic reported 83.1% for Claude Mythos Preview in April 2026, and OpenAI reported 85.6% for its updated GPT-5.5-Cyber. Sakana's score is a slight advancement beyond these figures. The CTI-REALM benchmark presents a different competitive landscape, with Microsoft's own evaluations indicating a distinct performance level for Fugu-Cyber.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.