By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities, most notably featuring a context window of 1 million tokens. This expanded context window allows the model to process and analyze vastly larger amounts of information in a single prompt, including entire books, hours of video, or extensive codebases. Previously, models like Gemini 1.0 Pro had a context window of 32,000 tokens, making the 1 million token capacity a more than 30-fold increase. This leap in context length is expected to unlock new applications in areas requiring deep understanding of lengthy documents or complex datasets.
The Gemini 1.5 Pro model is built on a new Mixture-of-Experts (MoE) architecture, which Google states makes it more efficient and performant than previous models. The MoE architecture allows the model to selectively activate different parts of its neural network for specific tasks, leading to faster processing and reduced computational cost. Google demonstrated the model's capabilities by having it analyze a 402-page PDF document, a 44-minute silent film, and 11 hours of audio, all within a single prompt. The model was able to recall specific details from these diverse inputs, showcasing its enhanced comprehension and recall abilities.
This development positions Gemini 1.5 Pro as a powerful tool for developers and businesses looking to leverage AI for complex analytical tasks. The expanded context window is particularly beneficial for tasks such as summarizing lengthy reports, analyzing legal documents, understanding intricate code, or even processing and cross-referencing information from multiple sources simultaneously. Google highlighted that the model can process over 1 million tokens in under 10 seconds, demonstrating its speed and efficiency despite the massive input size. The company is making Gemini 1.5 Pro available to developers via Google AI Studio and Vertex AI, with a tiered rollout of the 1 million token context window, starting with a smaller subset of users.
Google emphasized that while the 1 million token context window is the headline feature, Gemini 1.5 Pro also retains the multimodal reasoning capabilities of its predecessor, Gemini 1.5. This means it can understand and reason across different types of information, including text, images, audio, and video. The company is also working on making even larger context windows available in the future, signaling a continued push towards models that can handle increasingly vast amounts of data. The introduction of Gemini 1.5 Pro represents a significant step forward in the field of artificial intelligence, pushing the boundaries of what AI models can understand and process.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.