By Interestana AI Editorial — AI-drafted, human-overseen. How we report
AI Struggles With Ikea Furniture Assembly Benchmark

Researchers at Epoch AI have developed a novel benchmark, the Furniture Assembly Benchmark, designed to assess the actual intelligence of artificial intelligence models by testing their spatial reasoning and error detection capabilities. This benchmark utilizes the notoriously complex task of assembling Ikea furniture, a process often associated with relationship strain and significant problem-solving challenges. The researchers, Aiden Ament and Greg Burnham, conceptualized the benchmark during a company retreat, initially exploring methods where they would act as an AI's "robot" and follow literal instructions, which often led to "messy" results. They refined the approach to have AI models identify errors in partially assembled Ikea furniture, reflecting the practical reality that mid-assembly mistakes are acceptable if they can be caught and corrected. This method provides a more robust evaluation of an AI's intelligence compared to simply asking it for assembly instructions. The benchmark aims to offer a more concrete and verifiable way to compare the "smartness" of different AI models, moving beyond abstract claims of intelligence often made by tech companies. The researchers selected three Ikea products for their experiment, chosen based on their varying difficulty levels as indicated by Ikea's own Complexity Index. These included the Ställ shoe rack, designated as an "easy" level assembly, the Tonstad bed frame for a "medium" level challenge, and the Malm dresser, representing a "hard" level of complexity. By presenting AI models with images of partially assembled furniture, some containing deliberate errors, the researchers evaluated the models' ability to pinpoint these mistakes. While no AI model achieved a perfect score of 100% accuracy in identifying all errors across the tested furniture items, the results indicated that AI models are rapidly improving in this domain. The benchmark's focus on error identification is particularly significant, as it mimics real-world problem-solving scenarios where recognizing and rectifying mistakes is crucial for successful task completion. This approach offers a tangible, real-world application for evaluating AI's cognitive abilities, moving beyond theoretical benchmarks. The development of the Furniture Assembly Benchmark by Epoch AI, a nonprofit research organization focused on artificial intelligence, highlights the ongoing efforts within the AI community to create more objective and practical measures of artificial intelligence. The benchmark's design, centered on a universally recognized difficult task, makes it accessible for understanding and comparison, even as AI capabilities continue to advance at a rapid pace. The researchers' findings suggest that while current AI models demonstrate impressive progress in spatial reasoning and error detection, there remains a gap between their performance and human-level competence in complex, multi-step assembly tasks.
Original source — read the full reporting at the publisher:
Read on Fast CompanyGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.