By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Anthropic Expands AI Testing for Cyber Teams, Finds 129,000 Flaws

Anthropic announced on Tuesday, October 22, 2024, an expansion of its program that grants vetted cybersecurity professionals access to its advanced artificial intelligence models with diminished safeguards and blocking classifiers. This initiative, known as Project Glasswing, has reportedly identified a substantial number of software vulnerabilities. Between April 2026 and July 2026, the program uncovered at least 129,000 verified software vulnerabilities. The company further stated that it also identified an additional 20,000 potential vulnerabilities during the same period. These findings highlight the dual nature of advanced AI: its potential to aid in cybersecurity defense while also presenting risks that require rigorous testing and mitigation strategies. The expansion of this program suggests Anthropic's commitment to proactively addressing the security implications of its AI technologies by leveraging the expertise of the cybersecurity community. By allowing these professionals to probe the models with fewer restrictions, Anthropic aims to identify and rectify potential weaknesses before they can be exploited by malicious actors. The specific number of 129,000 verified vulnerabilities underscores the scale of the challenge in securing complex AI systems. These vulnerabilities could range from logical flaws that lead to incorrect outputs to more critical security loopholes. The additional 20,000 potential vulnerabilities indicate a broader scope of issues that require further investigation and validation. The program's focus on "vetted cybersecurity professionals" implies a controlled environment where participants are trusted entities, likely bound by strict non-disclosure agreements and ethical guidelines. This approach is crucial for responsible AI development, ensuring that discovered vulnerabilities are reported and addressed constructively rather than being leaked or misused. Anthropic's AI models, such as the Claude series, are known for their sophisticated natural language processing and reasoning capabilities. Making these models available for security testing allows researchers to explore how these advanced functionalities might be manipulated or exploited. The success of Project Glasswing in identifying such a large volume of flaws suggests that even highly advanced AI systems are not immune to security vulnerabilities. This reinforces the ongoing need for continuous security auditing and improvement in the AI domain. The company's proactive stance through Project Glasswing is a significant step in fostering a more secure AI ecosystem. It demonstrates a willingness to collaborate with external security experts to enhance the robustness and safety of its AI products. The findings from this project will likely inform future development cycles and security protocols for Anthropic's AI offerings, aiming to build more resilient and trustworthy AI systems for a wide range of applications.
Original source — read the full reporting at the publisher:
Read on The Hacker NewsGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.