By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities, most notably featuring an unprecedented 1 million token context window. This expanded context window allows the model to process and analyze vastly larger amounts of information in a single prompt, far exceeding the typical limits of previous models. For comparison, Gemini 1.5 Pro can ingest the equivalent of over 1,500 pages of text, or approximately one hour of video, or 11 hours of audio, or code from over 30,000 lines of code. This capability is a substantial leap from the 128,000 token context window offered by Gemini 1.0 Pro.
The enhanced context window is powered by Google's new Mixture-of-Experts (MoE) architecture, which enables more efficient processing of information. This architecture allows the model to selectively activate different parts of its neural network depending on the task, leading to faster inference times and reduced computational costs. Google demonstrated the model's capabilities by having it analyze a 44-minute silent film from 1927, identifying specific objects and events within the video, and answering detailed questions about its content. This showcases the model's potential for complex reasoning and analysis across various media types.
Gemini 1.5 Pro is also designed to be highly multimodal, meaning it can understand and process information from different formats, including text, images, audio, and video, simultaneously. This multimodal understanding is crucial for applications that require a holistic comprehension of complex data. The model is available in a preview for developers via the Google AI Studio and Vertex AI platforms, allowing them to experiment with its advanced features. Google has indicated that the 1 million token context window will be available to a limited set of developers initially, with broader availability planned for the future.
This development positions Gemini 1.5 Pro as a powerful tool for a wide range of applications, including advanced research, content summarization, code analysis, and complex problem-solving. The ability to process such extensive context efficiently could revolutionize how developers and researchers interact with and leverage AI for intricate tasks. Google's commitment to pushing the boundaries of AI context windows and multimodal processing underscores its ongoing efforts to lead in the artificial intelligence landscape.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.