By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities. This new model introduces a context window of 1 million tokens, a substantial increase from the standard 128,000 tokens offered by its predecessor, Gemini 1.0 Pro. This expanded context window allows Gemini 1.5 Pro to process and analyze vastly larger amounts of information in a single prompt, including entire codebases, lengthy books, or hours of video content. The model's architecture is based on a Mixture-of-Experts (MoE) approach, which Google states makes it more efficient and performant. The MoE architecture routes queries to specialized sub-networks within the larger model, optimizing resource utilization and inference speed. This efficiency is crucial for handling the computational demands of such a large context window. During internal testing, Gemini 1.5 Pro demonstrated the ability to recall specific details from a 402-page document and a 44-minute silent film, showcasing its capacity for deep comprehension and information retrieval across diverse media. The model also achieved a 99.9% accuracy rate in identifying specific lines of code within a 100,000-line codebase. This capability is particularly relevant for developers and researchers who need to analyze complex software projects. Google highlighted that the 1 million token context window is available in a private preview for developers, with plans for broader availability in the future. The company emphasized that this feature is designed to unlock new use cases for AI, such as summarizing lengthy legal documents, analyzing extensive research papers, or understanding complex financial reports. The development of Gemini 1.5 Pro represents a continued push by Google to enhance the reasoning and comprehension abilities of its AI models, positioning it as a key player in the competitive AI landscape. The increased context window is expected to enable more nuanced and sophisticated interactions with AI, allowing users to leverage AI for tasks that require understanding and processing large volumes of unstructured data. The model's multimodal capabilities, inherited from the Gemini family, also allow it to process and reason across text, images, audio, and video, further expanding its utility. The company stated that the model will be made available through Google AI Studio and Vertex AI, Google Cloud's enterprise AI platform, providing developers with tools to build and deploy AI-powered applications. The focus on developer access and enterprise solutions indicates Google's strategy to integrate its advanced AI models into practical business applications.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.