Interestana
Home/News/Cantina's apex-flash-1 Model Solves 40 of 60 Security Bug Tasks
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Cantina's apex-flash-1 Model Solves 40 of 60 Security Bug Tasks

Cantina Security, in collaboration with Yeta Labs, has introduced apex-flash-1, an open-weights artificial intelligence model specifically engineered for vulnerability research. This model is a reinforcement learning fine-tune of Z.ai's GLM-5.3-Flash and has been made available on Hugging Face under the permissive MIT license. The apex-flash-1 model is deployable using popular frameworks such as vLLM, SGLang, or Transformers, though its BF16 version requires approximately 640 GB of GPU memory. The model boasts a total of 312.3 billion parameters, as indicated in its Hugging Face safetensors metadata. Its base, GLM-5.3-Flash, is a Mixture-of-Experts model featuring 18 billion active parameters. Cantina employed a training methodology that included GRPO, a rank-256 LoRA, and selective full-parameter training. The training dataset comprised 150 tasks derived from 50 real-world vulnerability cases, with each case presented in three distinct variants: guided whitebox, focused whitebox, and focused blackbox. The vulnerability types covered include authorization, identity, and scope flaws, which constitute 72% of the cases. Additionally, accounting and numerical precision bugs account for 18% of the dataset, with the remaining cases involving time validation, business rules, and Server-Side Request Forgery (SSRF) vulnerabilities. According to the model card published on Hugging Face, the reinforcement learning rollouts were conducted within the Codex agent harness, utilizing environments that closely mimic production software and protocols. In an evaluation using 60 tasks drawn from 20 distinct, held-out vulnerability cases, apex-flash-1 demonstrated significant capabilities. The model successfully solved 40 out of these 60 tasks, achieving a pass@1 rate of 66.7%, at an estimated cost of $2.38 per run. For comparative analysis, the base GLM-5.3-Flash model solved 36 out of 60 tasks (60.0% pass@1) at a cost of approximately $4.56. When benchmarked against Anthropic's Claude 3 Opus 5 High, apex-flash-1's performance was competitive, with Opus solving 43 out of 60 tasks (71.7% pass@1) but at a substantially higher cost of about $74.68 per run. This indicates that apex-flash-1 achieved a solved task cost of roughly $0.06, compared to $1.74 for Claude Opus. These performance figures and costs are based on internal benchmarks reported by Cantina Security. Cantina positions apex-flash-1 not as a standalone orchestrator, but as a specialized worker model designed to be integrated into larger AI systems. Its intended applications include code reading, tool utilization, exploit development, and verification tasks. An experimental variant, apex-flash-1-abliterated, which features modified refusal behavior, has also been developed but was not separately evaluated. The company's rationale is that security defenders require highly capable AI models to effectively identify and address vulnerabilities.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next