Home/News/Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
The Economist3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities. This new iteration of Google's flagship AI model introduces a groundbreaking 1 million token context window, a tenfold increase over its predecessor, Gemini 1.0 Pro. This expanded context window allows the model to process and analyze vastly larger amounts of information in a single input, including entire codebases, lengthy books, or hours of video content. The enhanced capacity is achieved through a novel Mixture-of-Experts (MoE) architecture, which Google states is more efficient and performant than previous models.

During the announcement, Google demonstrated Gemini 1.5 Pro's ability to analyze a 402-page PDF document, a 44-minute YouTube video, and over 11 hours of audio, all within a single prompt. The model successfully identified specific details and answered complex questions based on the provided content, showcasing its improved comprehension and recall. This capability is particularly impactful for developers and researchers who can now leverage Gemini 1.5 Pro to understand and interact with extensive datasets without the limitations of smaller context windows. The model is also capable of multimodal reasoning, meaning it can process and understand information from different modalities, such as text, images, audio, and video, simultaneously.

Gemini 1.5 Pro is built on Google's latest AI advancements and is designed to be highly efficient, offering performance comparable to Gemini 1.0 Pro despite its significantly larger context window. Google has made the model available to developers and enterprise customers through the Google AI Studio and Vertex AI platforms. The company highlighted that while the 1 million token context window is available in preview, a standard version with a 128,000 token context window will be the default for most users. This strategic rollout allows for rigorous testing and feedback collection before a wider release of the full capacity.

The introduction of Gemini 1.5 Pro with its expansive context window positions Google at the forefront of AI development, addressing a key challenge in the field: the ability of models to retain and process information over extended periods. This breakthrough is expected to unlock new applications in areas such as code analysis, long-form content summarization, and complex data interpretation, potentially transforming how businesses and individuals interact with AI. The underlying MoE architecture is a key enabler of this performance leap, allowing for more specialized processing of different types of information within the model.

Original source — read the full reporting at the publisher:

Read on The Economist

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next