By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities. This new model introduces a context window of 1 million tokens, a substantial increase from previous models, allowing it to process and analyze vastly larger amounts of information in a single prompt. The expanded context window means Gemini 1.5 Pro can ingest and reason over entire codebases, lengthy books, or hours of video content, enabling more complex and nuanced understanding.
This capability was demonstrated by processing a 402-page PDF document, a 11-hour YouTube video, and over 10,000 lines of code. The model's performance was benchmarked against its predecessor, Gemini 1.0 Pro, showing a 50% improvement in reasoning tasks. Google highlighted that Gemini 1.5 Pro maintains high performance across various modalities, including text, image, audio, and video, despite the expanded context. The model is built on a Mixture-of-Experts (MoE) architecture, which Google states makes it more efficient and scalable for training and inference.
Gemini 1.5 Pro is currently available in a limited preview for developers and enterprise customers through the Google AI Studio and Vertex AI platforms. Google emphasized its commitment to responsible AI development, noting that safety filters have been rigorously tested and are being continuously improved. The company also plans to make the 1 million token context window available to a broader set of users in the coming months, with a gradual rollout expected. This development positions Gemini 1.5 Pro as a powerful tool for tasks requiring deep comprehension of extensive data sets, potentially transforming fields like research, software development, and content analysis.
The introduction of such a large context window addresses a key limitation in current AI models, which often struggle to retain information from long inputs. By enabling the model to 'remember' and reference more data, Gemini 1.5 Pro can perform more sophisticated analyses, such as summarizing lengthy reports, answering complex questions based on extensive documentation, or identifying subtle patterns across large datasets. The MoE architecture, a departure from the dense transformer models, allows specific parts of the neural network to be activated for different tasks, leading to more efficient computation and potentially faster response times for complex queries. Google's ongoing investment in AI research and development, exemplified by Gemini 1.5 Pro, underscores its competitive stance in the rapidly evolving artificial intelligence landscape.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.