By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced on February 15, 2024, that its Gemini 1.5 Pro large language model now supports a context window of 1 million tokens. This significant expansion allows the model to process and analyze vastly larger amounts of information in a single prompt compared to previous versions. The previous standard context window for many advanced models, including earlier iterations of Gemini, typically ranged from 32,000 to 128,000 tokens. A token is a piece of a word, and a 1 million token context window means Gemini 1.5 Pro can ingest and reason over the equivalent of hundreds of thousands of words, or roughly 1,500 pages of text, in one go. This capability is particularly impactful for tasks involving extensive documents, lengthy code repositories, or hours of video content.
This new context window is available in a preview version of Gemini 1.5 Pro, accessible to developers via the Google AI Studio and Vertex AI platforms. Google highlighted specific use cases, such as analyzing a 402-page PDF document, a 11-hour YouTube video, or over 15,000 lines of code. The model demonstrated its ability to quickly retrieve specific information and identify patterns within these large datasets. For instance, when presented with a 402-page document detailing the Apollo 11 mission, Gemini 1.5 Pro could accurately answer questions about specific technical details and mission objectives. Similarly, when provided with an 11-hour video, it could pinpoint specific moments and summarize events.
The underlying technology enabling this leap is a new Mixture-of-Experts (MoE) architecture that Google has implemented in Gemini 1.5 Pro. This architecture allows the model to be more efficient and scalable, activating only specific parts of the neural network relevant to the task at hand. This contrasts with traditional dense models that use all parameters for every computation. Google stated that this MoE approach is a key factor in achieving the massive context window while maintaining performance and reducing computational costs. The company also emphasized that the model's performance remains strong, with the standard 128,000 token context window version of Gemini 1.5 Pro outperforming Gemini 1.0 Pro across a range of benchmarks.
Google's move to a 1 million token context window positions Gemini 1.5 Pro as a leading model for complex information processing tasks. This capability is expected to accelerate development in areas such as legal document review, scientific research analysis, software development, and advanced content summarization. The availability of this feature in preview allows developers to experiment with and build applications that leverage this extended understanding. Google plans to make the 1 million token context window generally available in the coming months, following feedback and further refinement during the preview period. The company also indicated that future versions of Gemini may offer even larger context windows.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.