Interestana
Home/News/Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
The Economist3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities. The core innovation of Gemini 1.5 Pro is its massive 1 million token context window, a tenfold increase over the previous standard of 100,000 tokens. This expanded context window allows the model to process and reason over vastly larger amounts of information in a single prompt, including entire books, lengthy codebases, or hours of video content. For instance, users can input up to 1,100 pages of text, 11 hours of video, or 30,000 lines of code, enabling more comprehensive analysis and understanding.

This capability is powered by a new Mixture-of-Experts (MoE) architecture, which Google states makes Gemini 1.5 Pro more efficient and performant. The MoE approach allows the model to selectively activate specific parts of its neural network for different tasks, rather than engaging its entire capacity for every query. This optimization contributes to faster processing times and reduced computational costs, making advanced AI more accessible. Google has demonstrated this by having Gemini 1.5 Pro analyze a 402-page PDF document in under two seconds, showcasing its speed and efficiency in handling extensive data.

The model's multimodal capabilities are also enhanced, allowing it to understand and integrate information from various formats, including text, images, audio, and video. This holistic understanding is crucial for complex tasks that require synthesizing information from diverse sources. For developers, Google is offering Gemini 1.5 Pro through the Google AI Studio and Vertex AI platforms, providing tools to integrate its advanced features into their applications. The company is also making available a limited preview of Gemini 1.5 Flash, a lighter and faster version optimized for high-volume, lower-latency tasks, further expanding the utility of the Gemini family of models.

Google highlighted several use cases for the 1 million token context window, including summarizing lengthy research papers, analyzing complex legal documents, and understanding the narrative arc of entire movies. In a demonstration, Gemini 1.5 Pro was able to identify a specific object within a 48-minute silent film, showcasing its ability to recall and pinpoint details from extensive video inputs. The introduction of Gemini 1.5 Pro represents a significant step forward in AI's ability to comprehend and interact with large-scale, complex datasets, paving the way for more sophisticated applications across various industries.

Original source — read the full reporting at the publisher:

Read on The Economist

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next