Interestana
Home/News/Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
The Economist3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, featuring a groundbreaking 1 million token context window. This significant expansion allows the AI model to process substantially larger volumes of information, including entire codebases, multiple lengthy documents, or hours of video content, in a single prompt. The previous standard for many large language models hovered around 32,000 to 128,000 tokens, making the 1 million token capacity a substantial leap forward in AI's ability to understand and reason over extensive data.

This enhanced context window is powered by Google's new Mixture-of-Experts (MoE) architecture, which enables more efficient processing of information. The MoE approach allows the model to selectively activate different parts of its neural network for specific tasks, leading to improved performance and reduced computational overhead compared to traditional dense models. Gemini 1.5 Pro's ability to handle such a large context window means it can maintain coherence and recall details across much longer interactions or data sets. For developers, this translates to the potential for more sophisticated applications that require deep understanding of complex information, such as analyzing extensive legal documents, summarizing lengthy research papers, or debugging large software projects.

During its preview, Gemini 1.5 Pro demonstrated its capabilities by processing over 1 million tokens of text and code, and even analyzing a 402-page PDF document. The model also showcased its multimodal reasoning by processing a 44-minute silent film, identifying specific objects and events within it. This multimodal capability, combining text, image, audio, and video understanding, is a key differentiator for Gemini 1.5 Pro. Google highlighted that the model can recall specific details from the film, such as the brand of a car or the color of a character's shirt, demonstrating a high degree of accuracy and comprehension.

The introduction of Gemini 1.5 Pro follows Google's earlier release of Gemini 1.0 in December 2023, which established the foundation for its multimodal AI capabilities. The Gemini family of models is designed to be natively multimodal, meaning they can understand and operate across different types of information seamlessly. The 1.5 Pro version represents a significant upgrade in terms of scale and efficiency, making advanced AI tools more accessible and powerful for a wider range of applications. Google stated that the 1 million token context window will be available to developers through the Gemini API in Google AI Studio and Vertex AI, with a smaller 128,000 token window available by default for broader accessibility.

Original source — read the full reporting at the publisher:

Read on The Economist

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next