By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Anthropic Releases Claude Haiku 5.5 With 1M Context
Anthropic released Claude Haiku 5.5 this week, positioning it as its most affordable and rapid small-scale AI model to date. This new iteration is engineered for high-volume tasks such as text summarization, data compaction, classification, and the operation of sub-agents within larger AI systems. A key feature retained from previous versions is its substantial 1 million token context window, complemented by an output token capacity of up to 128,000 tokens. The pricing structure for Haiku 5.5 begins at $0.10 per million input tokens and $0.50 per million output tokens. Anthropic reports this represents a significant cost reduction, approximately 90% less than Claude Haiku 4.5 for prompts up to 100,000 tokens. The model is accessible via a hosted API and is generally available across multiple cloud platforms, including the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and the Claude Platform on AWS.
Claude Haiku 5.5 introduces an adjustable "effort" setting, a novel capability for the Haiku class of models. This "adaptive thinking" feature is enabled by default, with the effort parameter set to a medium level. The model processes both text and image inputs and generates text outputs, possessing a knowledge cutoff date of June 2026. For batch processing, the system supports up to 300,000 output tokens, currently in beta. Two important technical considerations for users migrating to Haiku 5.5 are noted: attempts to use non-default values for temperature, top_p, or top_k parameters will result in a 400 error. Additionally, the model's new tokenizer processes text such that it consumes approximately 30% more tokens compared to Haiku 4.5 for the same textual content; a migration guide is available to address these changes.
The pricing model for Haiku 5.5 employs a two-tier structure based on prompt token volume. For prompts containing up to 100,000 tokens, input costs are $0.10 per million tokens, and output costs are $0.50 per million tokens. Cache reads are priced at $0.01, and 5-minute cache writes are $0.125. For prompts exceeding 100,000 tokens, the rates increase to $0.50 per million for input and $2.50 per million for output. This contrasts with Haiku 4.5's pricing of $1 per million input and $5 per million output. Anthropic estimates that approximately 90% of Haiku 4.5 requests fell within the 100,000 token threshold. Factoring in the tokenizer adjustments, Anthropic projects that Haiku 5.5 will be, on average, about 75% cheaper for users. Further cost reductions of an additional 50% are achievable through batch processing. In comparison, GPT-6 Luna, another model in the market, lists identical short-context rates but initiates its higher pricing tier above 272,000 input tokens, at $0.20 for input and $0.75 for output. For a 150,000 token prompt, Luna is listed as cheaper.
Anthropic has also released benchmark performance data for Haiku 5.5, though these figures are self-reported. On the OSWorld 2.1 offline subset, Haiku 5.5 achieved a score of 72.4%. This performance is significantly higher than GPT-6 Luna's reported 48.9% and Haiku 4.5's 15.7%. In the Terminal-Bench 4.0 benchmark, Haiku 5.5 scored 39.2%, compared to Luna's 16.4% and Haiku 4.5's 0.0%. These benchmarks suggest a substantial leap in performance for Haiku 5.5, particularly in tasks represented by these specific evaluations, positioning it as a competitive option for developers seeking efficient and cost-effective AI solutions for large-scale text processing and analysis.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.