Home/News/AI Models Show Improved Reasoning on Complex Tasks
Delish4 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

AI Models Show Improved Reasoning on Complex Tasks

AI Models Show Improved Reasoning on Complex Tasks

Recent evaluations of artificial intelligence models indicate significant progress in their ability to handle complex reasoning and problem-solving tasks. These advancements are being tracked through a variety of new benchmarks designed to push the boundaries of current AI capabilities. One such benchmark, the "Complex Reasoning Challenge" (CRC), has shown that leading models are now achieving scores that were previously considered unattainable, suggesting a leap in their analytical and inferential capacities. The CRC evaluates models on their performance across diverse domains, including scientific hypothesis generation, logical deduction in intricate scenarios, and multi-step problem decomposition. Early results from the CRC, detailed in a preprint study released on May 15, 2024, by researchers at the AI Advancement Institute, show that models like "Cognito-7B" and "Nexus-12B" have surpassed previous state-of-the-art performance by an average of 18%. Cognito-7B, developed by the independent research collective "Synthetica Labs," is a 7-billion parameter model that has demonstrated particular strength in mathematical reasoning, solving 92% of advanced calculus problems presented in the benchmark. Nexus-12B, a 12-billion parameter model from "Global AI Corp," excels in logical reasoning and abstract problem-solving, achieving a 95% accuracy rate on tasks requiring the identification of subtle logical fallacies and the construction of coherent arguments. These improvements are attributed to novel training methodologies, including enhanced reinforcement learning from human feedback (RLHF) and the integration of "chain-of-thought" prompting techniques during the training phase, which encourage models to articulate their reasoning process step-by-step. The development of these more capable AI systems has implications across numerous sectors, from scientific research and drug discovery to financial modeling and advanced software development. For instance, in scientific research, AI models capable of complex reasoning could accelerate the process of identifying potential drug candidates or formulating new hypotheses by analyzing vast datasets and identifying non-obvious correlations. In finance, these models could lead to more sophisticated risk assessment tools and algorithmic trading strategies. The AI Advancement Institute's study also highlighted the emergence of "emergent abilities" in these larger models, where capabilities not explicitly trained for appear as model size and complexity increase. This phenomenon suggests that current benchmarks may not fully capture the full spectrum of potential AI advancements. However, the researchers also caution that these models still exhibit limitations, particularly in areas requiring common sense reasoning and understanding of nuanced social contexts. The benchmark results are expected to spur further research into developing more robust and generalizable AI systems. Global AI Corp announced on May 16, 2024, that they are already working on a successor to Nexus-12B, aiming to incorporate more sophisticated contextual understanding. Synthetica Labs has also indicated plans to release a more accessible version of Cognito-7B for academic research by the end of the third quarter of 2024. The ongoing development and evaluation of these advanced AI models underscore the rapid pace of innovation in the field and the increasing need for comprehensive and challenging assessment tools.

Original source — read the full reporting at the publisher:

Read on Delish

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next