By Interestana AI Editorial — AI-drafted, human-overseen. How we report
AI Models Struggle With Intelligence Tests
Artificial intelligence models are demonstrating significant progress in solving complex puzzles, yet they continue to struggle with certain types of intelligence tests, revealing persistent cognitive differences compared to human reasoning. Puzzles and games have historically served as crucial benchmarks for AI development, dating back to Arthur Samuel's work on a checkers-playing algorithm in 1959, which popularized the term "machine learning." More recently, AI has excelled at tasks like the New York Times Connections puzzles; a Columbia University study in late 2024 showed models solving only 18% of these puzzles, a figure that improved to near-perfect scores by early 2025. This rapid advancement underscores AI's growing capabilities, but the areas where models still falter offer valuable insights into the technology's limitations and the unique strengths of human cognition. Despite considerable progress, current AI models often fail when faced with subtle alterations in classic riddles and exhibit particular weaknesses in visual puzzles. These challenges provide opportunities to test human intelligence against that of AI, with some puzzles proving as difficult for humans as they were for AI at one point, while others highlight the perceived simplicity of human problem-solving in contrast to AI's current limitations. The divergence in performance on these tests emphasizes the distinct nature of machine and human cognitive processes. Spatial reasoning, a domain where humans typically possess a significant advantage, presents a notable challenge for AI. Mental rotation problems, common in IQ tests, require individuals to identify if different images depict the same object from varying perspectives. Although contemporary language models are increasingly capable of processing visual inputs, they continue to perform poorly on these spatial reasoning tasks. This persistent difficulty in visual-spatial tasks, alongside their struggles with nuanced linguistic puzzles, indicates that while AI can master pattern recognition and data processing at an unprecedented scale, it has yet to replicate the intuitive and flexible reasoning that characterizes human intelligence in these specific areas. The ongoing development of AI aims to bridge these gaps, but for now, human performance on these select intelligence tests remains superior.
Original source — read the full reporting at the publisher:
Read on MIT Technology ReviewGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.