By Interestana AI Editorial — AI-drafted, human-overseen. How we report
AI Self-Improvement May Take Longer Than Predicted
The prospect of artificial intelligence achieving recursive self-improvement, where AI systems autonomously enhance their own capabilities without significant human intervention, may be further off than anticipated, according to a recent study. While current large language models (LLMs) demonstrate proficiency in tasks such as code generation, synthetic data creation, and optimizing hardware, a multi-institution research group has found that AI agents are not yet equipped for the open-ended, judgment-driven nature of AI research itself. This finding challenges optimistic forecasts that predict rapid automation of AI development.
The study, led by Peter Kirgis and Sayash Kapoor at Princeton University, evaluated AI agents' capacity for conducting AI research. The researchers concluded that while AI can effectively address the engineering challenges inherent in research, it currently lacks the crucial judgment and creativity needed to produce original research at the level required for acceptance at top machine-learning conferences. This gap indicates that some projected timelines for automating AI research might be overly ambitious, not yet fully supported by current AI capabilities.
Existing evaluations of AI agents' research automation potential often focus on their ability to complete narrow tasks with verifiable outcomes, such as solving specific engineering problems or fine-tuning small language models against predefined benchmarks. However, genuine progress in AI research necessitates open-ended thinking. This includes formulating hypotheses, determining the evidence required to validate them, and recognizing when to abandon an approach and restart. These are complex cognitive processes that current AI agents struggle to replicate independently.
To assess these higher-order research skills, the researchers introduced a novel evaluation method termed "shadow evaluation." This technique requires an AI agent to address a research question drawn from a high-quality, unpublished academic paper. In their experiment, the researchers tasked Anthropic's Claude Opus 4.8, operating via the open-source software OpenClaw, with answering such questions. The specific questions were sourced from two papers submitted to the NeurIPS 2026 conference, a prestigious event in the machine-learning field. The initial question presented to the AI agent was designed to test its ability to engage with complex, unresolved research problems, moving beyond simple problem-solving to encompass the nuanced decision-making characteristic of scientific inquiry. This methodology aims to provide a more realistic assessment of AI's potential to contribute to the frontier of AI research, rather than just executing well-defined tasks.
Original source — read the full reporting at the publisher:
Read on MIT Technology ReviewGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.