By Interestana AI Editorial — AI-drafted, human-overseen. How we report
AI Model Claude 3.5 Sonnet Outperforms GPT-4o

Anthropic announced the release of its Claude 3.5 Sonnet AI model on June 20, 2024, positioning it as a significant advancement in large language model capabilities. The company stated that Claude 3.5 Sonnet outperforms all current open-source models and rivals or surpasses proprietary models like OpenAI's GPT-4o on a variety of industry benchmarks. This release marks Anthropic's continued effort to compete at the forefront of AI development.
Claude 3.5 Sonnet achieved a score of 92.2% on the MMLU (Massive Multitask Language Understanding) benchmark, a measure of general knowledge and problem-solving across 57 subjects. This score is higher than the 90.8% achieved by GPT-4o, according to Anthropic's internal evaluations. The model also demonstrated a 64% accuracy on the GPQA (Graduate-Level Google-Proof Questions) benchmark, exceeding GPT-4o's 61.5%. Furthermore, Claude 3.5 Sonnet showed a 71.9% accuracy on the MATH benchmark, surpassing GPT-4o's 68.4% and Meta's Llama 3 70B's 66.8%. These benchmarks are critical for assessing the reasoning and comprehension abilities of AI models.
Beyond raw benchmark performance, Anthropic highlighted Claude 3.5 Sonnet's enhanced capabilities in areas such as coding, mathematics, and nuanced instruction following. The model is designed to be faster and more cost-effective than its predecessors, making advanced AI more accessible. A key new feature introduced with this release is "Artifacts," which allows users to interact with and edit AI-generated content, such as code or text documents, directly within the Claude interface. This feature aims to streamline workflows and improve user productivity.
The introduction of Claude 3.5 Sonnet follows Anthropic's previous releases, including Claude 3 Opus, Sonnet, and Haiku, which were launched in March 2024. Opus was positioned as the most powerful model, Sonnet as the balance of intelligence and speed, and Haiku as the fastest. The 3.5 iteration of Sonnet represents a substantial leap forward, particularly in its ability to handle complex reasoning tasks and generate more accurate, contextually relevant outputs. The company emphasized that Claude 3.5 Sonnet is the first model to be released under the new 3.5 series, with further updates and model tiers expected.
Anthropic, founded by former OpenAI researchers, has been a prominent player in the AI race, focusing on safety and ethical AI development. The company's competitive releases aim to challenge the dominance of other major AI labs. The performance gains demonstrated by Claude 3.5 Sonnet suggest that the competition in the advanced AI model market remains intense, with significant innovation occurring rapidly. The availability of Claude 3.5 Sonnet is expected to influence the development and deployment of AI applications across various industries.
Original source — read the full reporting at the publisher:
Read on BBC SportGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.