Interestana
Home/News/Exa Launches Agent Ultra API for Exhaustive Research
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Exa Launches Agent Ultra API for Exhaustive Research

Exa has launched Agent Ultra, the most advanced tier of its Exa Agent API, engineered for research tasks requiring exhaustive data collection, such as extensive list building, entity enrichment, and answering complex questions that necessitate analysis of thousands of sources. The company reports that Agent Ultra surpasses Opus 5.5, GPT-6 Astra, and Perplexity Agent when each competitor is run at its maximum effort setting across four distinct research benchmarks. Agent Ultra is available today as a hosted API through the Exa API by specifying the "ultra" effort level. It is not an open-weights model and cannot be self-hosted by users. The Exa Agent system functions by dividing a research task into smaller subtasks, which are then assigned to specialized subagents that can investigate multiple domains concurrently. This architecture allows for the dynamic routing of advanced frontier models to steps requiring their capabilities, while employing faster models for less demanding stages. Agent Ultra represents the highest compute-intensive mode, designed to run for extended durations to ensure the most comprehensive results possible. According to the Agent Ultra documentation, complex tasks typically conclude within approximately 30 minutes, though exceptionally difficult tasks may extend up to 3 hours. Exa's launch post details benchmark results comparing Agent Ultra against competitors, all of which were also tested at their maximum effort settings. In the WANDR benchmark, which measures soft recall, Agent Ultra achieved an 81.4% score, significantly outperforming Opus 5.5 (72.3%), GPT-6 Astra (26.0%), and Perplexity Agent (40.1%). For the DeepSearchQA benchmark, measuring F1 score, Agent Ultra reached 93.9%, compared to Opus 5.5 (77.6%), GPT-6 Astra (85.3%), and Perplexity Agent (89.7%). The WideSearch benchmark, assessing row-level F1, saw Agent Ultra score 58.9%, ahead of Opus 5.5 (51.6%), GPT-6 Astra (54.7%), and Perplexity Agent (56.0%). In the Company Find-All benchmark, which counts the average number of passing entities per task, Agent Ultra identified 2,451 entities, a substantial increase over Opus 5.5 (146), GPT-6 Astra (113), and Perplexity Agent (98). Exa also provided cost-effectiveness comparisons. For the WANDR benchmark, Agent Ultra offered a 12.6% improvement over Opus 5.5 at half the cost per task. In DeepSearchQA, it achieved a 4.7% higher score than Perplexity Agent while costing 46% less per task than GPT-6 Astra. The WideSearch benchmark showed Agent Ultra performing 5.2% better than Perplexity Agent and incurring the lowest cost per task among the four systems. For Company Find-All, Agent Ultra demonstrated a 1579% increase in entities found compared to Opus 5.5, with the lowest cost per entity. These reported gains are relative improvements, not percentage point differences; for instance, the absolute gap between Agent Ultra and Opus 5.5 on WANDR is 9.1 points.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next