Interestana
Home/News/Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
The Economist3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in its artificial intelligence model capabilities. The core innovation of Gemini 1.5 Pro is its unprecedented 1 million token context window, a substantial increase from the 128,000 tokens available in Gemini 1.0 Pro. This expanded context window allows the model to process and analyze vastly larger amounts of information in a single input. For instance, it can ingest an entire book, a 400-page document, or over an hour of video content. This capability is powered by a new Mixture-of-Experts (MoE) architecture, which Google states makes the model more efficient and performant. The MoE architecture allows the model to selectively activate different parts of its neural network for specific tasks, rather than engaging the entire network for every operation. This approach is designed to improve processing speed and reduce computational costs. The 1 million token context window is currently available in a limited preview for developers and enterprise customers, with plans for broader availability in the future. Google highlighted several potential applications for this enhanced context processing. Developers can use it to build more sophisticated applications that require understanding long-form content, such as summarizing lengthy legal documents, analyzing complex codebases, or extracting information from extensive research papers. In the realm of video analysis, Gemini 1.5 Pro can process up to an hour of video, enabling tasks like identifying specific moments within a film or transcribing dialogue from extended footage. This represents a leap forward in multimodal AI, as the model can now more effectively integrate and reason across different types of data, including text, images, audio, and video. The development of Gemini 1.5 Pro follows the initial release of Gemini models in December 2023, which were designed from the ground up to be multimodal. The Gemini family of models includes Gemini Ultra, Gemini Pro, and Gemini Nano, each tailored for different scales of deployment, from data centers to on-device applications. Gemini 1.5 Pro, as a refined version of Gemini Pro, aims to offer a balance of advanced capabilities and accessibility for developers. Google's ongoing investment in AI research and development, particularly in large language models and multimodal AI, positions it as a key player in the rapidly evolving AI landscape. The company's strategy involves making these advanced AI capabilities available through its cloud platform, Google Cloud, to empower businesses and researchers to innovate. The introduction of such a large context window is expected to spur new research and development in AI applications that were previously constrained by the limitations of shorter context lengths.

Original source — read the full reporting at the publisher:

Read on The Economist

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next