By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the availability of Gemini 1.5 Pro, a new iteration of its flagship large language model, on February 15, 2024. This advanced model introduces a groundbreaking 1 million token context window, a significant leap from previous models that typically handled tens of thousands of tokens. This expanded context window allows Gemini 1.5 Pro to process and reason over vastly larger amounts of information, including entire codebases, lengthy books, or hours of video content, in a single prompt.
The 1 million token context window represents a substantial increase in the model's capacity for understanding and retaining information. For comparison, earlier models like Gemini 1.0 Pro had a context window of 32,000 tokens. This enhancement enables developers and users to input and analyze much more complex and extensive data sets. For instance, the model can now ingest up to 1,500 pages of text, 11 hours of video, or 30,000 lines of code. Google demonstrated this capability by having Gemini 1.5 Pro analyze a 402-page PDF document and a 44-minute silent film, identifying specific objects and moments within the film with remarkable accuracy.
Gemini 1.5 Pro is built on a new Mixture-of-Experts (MoE) architecture, which Google states makes it more efficient and performant. This architectural shift allows the model to selectively activate different parts of its neural network for specific tasks, leading to faster processing and reduced computational overhead compared to traditional dense models. The MoE architecture is a key factor in enabling the model to handle the immense context window while maintaining speed and accuracy. Google has made Gemini 1.5 Pro available in a limited preview for developers, with plans for broader access in the future.
The implications of such a large context window are far-reaching. It could revolutionize how AI models are used in various fields, from software development and legal analysis to scientific research and content creation. Developers can now build applications that require a deep understanding of extensive documentation or historical data. For example, a developer could feed an entire project's codebase into Gemini 1.5 Pro to identify bugs or suggest optimizations. Similarly, researchers could use it to analyze vast quantities of scientific literature or experimental data. The ability to process and recall information from such large inputs marks a significant step towards more capable and versatile AI systems.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.