Interestana
Home/News/Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
The Economist3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, featuring a groundbreaking 1 million token context window. This significant expansion allows the AI model to process and reason over substantially larger amounts of information than its predecessors, potentially revolutionizing how AI interacts with complex data sets. The previous standard for many large language models was around 128,000 tokens, making the 1 million token capacity a tenfold increase.

This enhanced context window means Gemini 1.5 Pro can ingest and analyze entire codebases, lengthy books, or hours of video content in a single prompt. For developers and researchers, this capability opens new avenues for tasks such as summarizing extensive documents, analyzing complex code for bugs or vulnerabilities, and extracting insights from large video archives. The model's ability to maintain performance and accuracy with such a large context is a key technical achievement, addressing a long-standing challenge in AI development where performance often degrades with increased input size.

Gemini 1.5 Pro is built on Google's latest AI architecture, incorporating a Mixture-of-Experts (MoE) model. This architectural choice contributes to its efficiency and scalability, allowing it to handle the demands of a 1 million token context window. The MoE approach enables the model to selectively activate different parts of its neural network for specific tasks, leading to faster processing and reduced computational overhead compared to monolithic models of similar scale. Google highlighted that the model can process up to 1 million tokens in under a minute, a feat that underscores its advanced engineering.

In addition to the expanded context window, Gemini 1.5 Pro demonstrates strong performance across a range of benchmarks, including multimodal reasoning. The model can understand and process information from various modalities, such as text, images, audio, and video, simultaneously. This multimodal capability is crucial for real-world applications where data is rarely confined to a single format. Google provided demonstrations showing Gemini 1.5 Pro analyzing a 402-page PDF document and a 44-minute silent film, extracting specific details and answering complex questions about their content. The public preview is available through the Google AI Studio and Vertex AI platforms, allowing developers to integrate this advanced AI into their applications and workflows.

Original source — read the full reporting at the publisher:

Read on The Economist

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next