By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities, most notably featuring a context window of 1 million tokens. This expanded context window allows the model to process and reason over substantially larger amounts of information than previous versions, including entire codebases, multiple lengthy documents, or hours of video. The 1 million token context window represents a tenfold increase over the 100,000 token limit of Gemini 1.0 Pro, positioning Gemini 1.5 Pro as a leader in handling extensive data inputs. The model's ability to process such a large context is achieved through a new Mixture-of-Experts (MoE) architecture, which Google states is more efficient and performant. This architecture allows the model to selectively activate relevant parts of its neural network for specific tasks, leading to faster processing and reduced computational overhead. Gemini 1.5 Pro's multimodal capabilities are also enhanced, enabling it to understand and analyze various data formats, including text, images, audio, and video, within this massive context. For instance, users can input entire books, extensive code repositories, or hours of video footage and ask complex questions or request summaries and analyses. Google demonstrated this capability by having Gemini 1.5 Pro analyze a 402-page PDF document and a 44-minute silent film, extracting specific details and answering complex queries about their content. The company highlighted that the model could recall specific details from the film, such as the exact moment a particular object appeared. The development of Gemini 1.5 Pro builds upon Google's ongoing efforts in AI research and development, aiming to create more powerful and versatile AI models. This release follows the initial launch of Gemini, which was introduced in December 2023 as Google's most capable AI model family to date, designed to be multimodal from the ground up. The 1 million token context window is currently available in a limited preview for developers and enterprise customers, with plans for broader availability in the future. Google emphasized its commitment to responsible AI development, stating that safety and ethical considerations are paramount in the deployment of these advanced models. The company also noted that while the context window is 1 million tokens, they are experimenting with even larger windows, suggesting future iterations could handle even more data. This advancement is expected to unlock new applications in areas such as complex data analysis, code comprehension, long-form content summarization, and advanced video understanding, potentially transforming how businesses and researchers interact with and leverage large datasets.
Original source — read the full reporting at the publisher:
Read on The EconomistGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.