Interestana
Home/News/Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Bon Appétit3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google announced the release of Gemini 1.5 Pro, an updated version of its flagship large language model, on February 15, 2024. The most significant advancement in this iteration is its expanded context window, now capable of processing up to 1 million tokens. This represents a substantial leap from the 128,000 token limit of Gemini 1.0 Pro, allowing the model to analyze vastly larger amounts of information in a single prompt. For instance, users can now input an entire book, a lengthy codebase, or over an hour of video content for analysis. The model also features a Mixture-of-Experts (MoE) architecture, which Google states improves efficiency and performance by activating only relevant parts of the neural network for specific tasks. This architecture is a key component in enabling the model's enhanced capabilities while managing computational resources. Gemini 1.5 Pro demonstrates a significant improvement in reasoning capabilities, particularly in understanding long, complex documents and code. Google highlighted its ability to perform tasks such as summarizing lengthy reports, extracting specific information from extensive datasets, and identifying patterns or anomalies within large code repositories. The model's performance in these areas is attributed to its enhanced understanding of context and its capacity to retain information over extended sequences. The company also emphasized that Gemini 1.5 Pro maintains the multimodal capabilities of its predecessor, meaning it can process and understand various types of information, including text, images, audio, and video. This multimodal understanding, combined with the massive context window, opens up new possibilities for complex analytical tasks. For example, it can analyze video content frame-by-frame or process audio transcripts alongside accompanying text to provide comprehensive insights. Google has made Gemini 1.5 Pro available in a limited preview for developers and enterprise customers, with plans for broader availability in the future. The company is also offering a version with a 128,000 token context window for general availability through the Gemini API. This phased rollout allows Google to gather feedback and refine the model based on real-world usage before a wider release. The development of Gemini 1.5 Pro positions Google at the forefront of advancements in large language models, particularly in the area of long-context understanding, which is crucial for many advanced AI applications. The expanded context window is expected to significantly impact fields such as legal document review, scientific research, software development, and media analysis, where processing large volumes of data is a common requirement. The company stated that the model achieved top performance on various benchmarks, including the MMLU (Massive Multitask Language Understanding) benchmark, with a score of 90.6%. This benchmark assesses a model's knowledge and reasoning abilities across 57 diverse subjects. The Gemini 1.5 Pro model also showed strong performance in coding benchmarks, demonstrating its utility for developers. Google's commitment to responsible AI development is also highlighted, with the company stating that safety and ethical considerations have been integrated throughout the model's training and evaluation process. The introduction of Gemini 1.5 Pro with its 1 million token context window marks a significant milestone in the evolution of AI, enabling more sophisticated and comprehensive data analysis than previously possible.

Original source — read the full reporting at the publisher:

Read on Bon Appétit

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next