By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the public release of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities. The core innovation of Gemini 1.5 Pro is its unprecedented 1 million token context window, a substantial increase from the typical few thousand tokens found in previous models. This expanded context window allows the AI to process and analyze vastly larger amounts of information in a single prompt, including entire books, lengthy codebases, or hours of video. For instance, a 1 million token context window can accommodate approximately 1,500 pages of text or 11 hours of video. This capability is crucial for complex tasks requiring a deep understanding of extensive data, such as summarizing lengthy documents, analyzing intricate legal texts, or comprehending the narrative arc of a full-length film. The model also demonstrates enhanced performance in multimodal reasoning, building upon the capabilities of its predecessor. Gemini 1.5 Pro is built on a Mixture-of-Experts (MoE) architecture, which Google states makes it more efficient and faster than standard dense models, particularly for long context tasks. This architecture allows the model to selectively activate different parts of its neural network for specific queries, optimizing computational resources. The initial rollout of Gemini 1.5 Pro is available through the Google AI Studio and Vertex AI platforms, targeting developers and enterprise users. Google highlighted several use cases, including analyzing a 402-page PDF document about the Apollo 11 mission, processing a 44-minute silent film to identify specific objects and actions, and analyzing a 128,000-line codebase. The company also noted that Gemini 1.5 Pro achieved top performance on several industry benchmarks, including MMLU (Massive Multitask Language Understanding) and HumanEval, which tests coding capabilities. This release positions Google to compete more directly with other leading AI providers that have been expanding their models' context windows and multimodal functionalities. The increased context length is expected to unlock new applications across various industries, from scientific research and education to entertainment and software development, by enabling AI to grasp and reason over much larger datasets than previously feasible. Google plans to make Gemini 1.5 Pro generally available later this year, with further enhancements and broader accessibility anticipated.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.