Interestana
Home/News/Children Outperform AI in Language Learning Efficiency
MIT Technology Review3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Children Outperform AI in Language Learning Efficiency

Children Outperform AI in Language Learning Efficiency

Human children achieve perfect fluency in human language with an efficiency that current artificial intelligence models cannot match, despite significant advancements in large language models (LLMs). For at least 100,000 years, children have been the sole entities capable of mastering language, a feat now shared by AI systems such as OpenAI's GPT models, Claude, and DeepSeek. These LLMs can converse naturally and convincingly, leading many to take their capabilities for granted.

However, the underlying process of teaching AI language reveals a substantial disparity known as the data efficiency gap. LLMs require an "inhuman amount of data" to achieve linguistic proficiency, consuming hundreds of thousands of times more words than a child experiences while mastering their native tongue. Michael C. Frank, a cognitive scientist at Stanford University, highlights this by stating that current AI training necessitates "burn[ing] down a forest and scrap[ing] the entire sum of all human knowledge" to replicate a milestone that occurs naturally in children over a year.

This gap presents a significant question for cognitive scientists and a challenge for AI developers: how do children achieve such superior linguistic learning efficiency compared to even the most sophisticated AI models? Addressing this question holds implications for both AI research and the field of cognitive science. Historically, the advancement of language models has primarily involved increasing their size and the volume of data they process. For instance, Meta's Llama 3.1, an open-weight LLM released two years ago, underwent pretraining on 15 trillion tokens, which are word-like units of language. Frontier models are reportedly pretrained on datasets ten times larger than this, according to Ethan Gotlieb Wilcox. The sheer scale of data used by these models underscores the efficiency advantage children possess, as they acquire language with far less exposure.

The data efficiency gap is a critical area of research, prompting investigations into the cognitive processes that enable children to learn language so effectively. Understanding these mechanisms could lead to the development of more efficient AI training methods, reducing the computational resources and data requirements for building advanced language models. Furthermore, it offers insights into the fundamental nature of human language acquisition and cognition. The ongoing efforts to bridge this gap involve exploring novel AI architectures and learning paradigms that mimic the human learning process more closely, moving beyond brute-force data scaling.

Original source — read the full reporting at the publisher:

Read on MIT Technology Review

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next