By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Anthropic Releases Claude Sonnet 5.5 With 70.6% Terminal-Bench 4.0 Score
Anthropic released Claude Sonnet 5.5 on June 26, 2024, positioning it as a faster and more cost-effective complement to its higher-tier Claude Opus 5.5 model. This new iteration is designed for well-scoped everyday tasks, bug fixing, and the creation of polished documents, slides, and spreadsheets. Claude Sonnet 5.5 is accessible via the Claude Platform under the identifier claude-sonnet-5-5, and is also available through major cloud providers including AWS, Google Cloud, and Microsoft Azure. As a closed-weights model, it does not support self-hosting.
Anthropic reports four key upgrades compared to its predecessor, Sonnet 5. Output generation speed is over 30% faster, making it the quickest Sonnet model to date. The cost per task is reduced by up to 30%, attributed to fewer token usages and tool calls. Improvements in writing capabilities are noted, with early testers describing it as a better collaboration partner and producing clearer prose. The model also demonstrates enhanced vision and long-horizon capabilities, notably becoming the first Sonnet model to successfully beat the game Pokémon Red using only screenshots.
Claude Sonnet 5.5 features a 1 million token context window and a maximum output of 128,000 tokens, with a knowledge cutoff date of June 2026. Adaptive thinking is enabled by default and can be adjusted across five effort levels: low, medium, high, xhigh, and max. In benchmark testing, Claude Sonnet 5.5 achieved a score of 70.6% on Terminal-Bench 4.0, a significant improvement from Sonnet 5's 10.3% and surpassing Opus 5.5's score of 66.4% at the Xhigh setting. On CursorBench 4.0, it scored 55.5%, slightly below Opus 5.5's 57.8%. FrontierCode 1.1 benchmarks show 52.1% at Xhigh and 46.2% at Max, with GPT-6 Sol scoring 49.3%. The GDPval-AA v2.1 benchmark registered 1844, very close to Opus 5.5's 1846 and higher than Sonnet 5's 1449. OSWorld 2.1 saw Sonnet 5.5 achieve 80.1% on computer use, nearing Opus 5.5's 81.8%. Humanity's Last Exam performance improved to 64.5% with tools, up from 54.9% in Sonnet 5. Anthropic clarified that the lower score at the 'Max' setting for FrontierCode was due to the model more frequently engaging in multi-agent code reviews, which sometimes led to timeouts or out-of-scope edits that penalized its score. The company also reiterated that Claude Opus 5.5 remains superior for complex, open-ended tasks.
Pricing for Claude Sonnet 5.5 remains unchanged from its predecessor, with input priced at $2 per million tokens and output at $10 per million tokens. This pricing structure aims to make the model more accessible for a wider range of applications, balancing performance gains with cost efficiency. Anthropic continues to emphasize the development of AI systems that are helpful, honest, and harmless, with ongoing research into safety and ethical considerations guiding their product releases.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.