Interestana
Home/News/Startups Seek Next Frontier Beyond Transformer LLMs
MIT Technology Review3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Startups Seek Next Frontier Beyond Transformer LLMs

A new wave of startups is actively seeking to advance the field of artificial intelligence by developing large language models (LLMs) that move beyond the foundational transformer architecture. This pursuit is driven by the recognition that while transformers, introduced in a 2017 Google paper titled "Attention Is All You Need," have powered the current generation of LLMs, they are beginning to exhibit limitations. These limitations are prompting researchers and engineers to explore novel approaches to LLM development. The transformer architecture's core strength lies in its "dense attention" mechanism, which encodes text meaning by comparing every word or token with every other word through multiplication. This process allows for remarkable accuracy in capturing textual meaning. However, this dense attention mechanism also presents significant challenges, particularly concerning computational efficiency and scalability as model sizes and input data increase. Many recent breakthroughs in LLMs, such as enhanced reasoning capabilities and the ability to process larger input contexts, are not direct evolutions of the transformer's core technology but rather workarounds designed to mitigate its inherent flaws. These workarounds highlight the growing need for fundamentally new architectures that can overcome the scaling and efficiency issues associated with dense attention. The current AI industry, which is largely built upon transformers, is now at a critical juncture. Justin Dangel, co-founder and CEO of the AI startup Subquadratic, emphasized the transformative impact of transformers, stating, "The entire AI industry is built on transformers. They are one of the most important innovations in the history of computer science, and they’ve changed the world." Despite their success, the limitations are becoming increasingly apparent, leading to a growing demand for innovation. Startups entering this space are aiming to push the boundaries of LLM technology, exploring new paradigms that could define the next generation of AI. These emerging companies, while facing the inherent risks of the startup ecosystem, have the potential to introduce disruptive innovations. They are less constrained by existing infrastructure and legacy systems compared to established players, allowing them to experiment more freely with novel architectural designs and training methodologies. The focus is on developing models that are not only more capable but also more efficient and scalable, addressing the computational bottlenecks that currently limit the advancement of LLMs. This exploration into post-transformer architectures signifies a potential paradigm shift in how LLMs are conceived and built, moving beyond incremental improvements to the existing transformer framework and towards entirely new computational approaches for natural language processing and understanding.

Original source — read the full reporting at the publisher:

Read on MIT Technology Review

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next