By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the release of Gemini 1.5 Pro, a significant advancement in its artificial intelligence model capabilities, on February 15, 2024. The core innovation of Gemini 1.5 Pro is its unprecedented 1 million token context window. This expanded context window allows the model to process and analyze vastly larger amounts of information in a single input compared to previous models. For reference, a token is a piece of a word, and a 1 million token context window can accommodate approximately 1,500 pages of text, or up to one hour of video, or over 30,000 lines of code. This capability dramatically enhances the model's ability to understand complex documents, long conversations, and extensive codebases.
This new model builds upon the foundation of Gemini 1.0, which was released in December 2023 and is Google's most capable AI model family to date. Gemini 1.5 Pro is designed to be multimodal, meaning it can understand and operate across different types of information, including text, images, audio, and video. The expanded context window is particularly impactful for video and audio analysis, enabling the AI to recall specific details from lengthy media files. For instance, a user could ask Gemini 1.5 Pro to find a specific moment in a 45-minute lecture video or identify a particular sound within an hour-long audio recording.
Google highlighted several use cases for this enhanced context window. Developers can leverage it to analyze entire code repositories to identify bugs or suggest optimizations. Researchers can feed lengthy scientific papers or historical documents into the model for comprehensive analysis and summarization. Business professionals can use it to process extensive market research reports or lengthy legal documents. The model's ability to maintain performance and accuracy even with such large inputs is a key differentiator. Google stated that Gemini 1.5 Pro demonstrates near-human performance on standard benchmarks, with a particular emphasis on its reasoning capabilities across modalities.
Gemini 1.5 Pro is currently available in preview for developers via the Google AI Studio and Vertex AI platforms. Google plans to make it generally available later this year. The company also announced that Gemini 1.5 Pro will offer enhanced performance and efficiency, with a smaller, more efficient Mixture-of-Experts (MoE) architecture that enables faster processing and reduced computational costs. This architecture allows the model to selectively activate different parts of its neural network for specific tasks, optimizing resource utilization. The development of Gemini 1.5 Pro represents a significant step forward in making AI models more practical and powerful for a wide range of real-world applications, pushing the boundaries of what AI can understand and process.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.