Interestana
Home/News/Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window
The Economist3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Google Unveils Gemini 1.5 Pro With 1 Million Token Context Window

Google announced the public preview of Gemini 1.5 Pro on February 15, 2024, a significant advancement in large language model capabilities. The core innovation of Gemini 1.5 Pro is its unprecedented context window, which has been expanded to 1 million tokens. This massive context window allows the model to process and analyze vastly larger amounts of information in a single prompt compared to previous models. For context, a token is roughly equivalent to four characters or about 0.75 words in English. Therefore, a 1 million token context window enables Gemini 1.5 Pro to ingest and reason over the equivalent of hundreds of thousands of words, or even entire books, at once.

This expanded context window is a breakthrough for several applications. It means Gemini 1.5 Pro can analyze lengthy documents, extensive codebases, hours of video, or hours of audio without losing track of details or requiring complex summarization techniques beforehand. Google demonstrated this capability by showing Gemini 1.5 Pro processing a 402-page PDF document and a 44-minute silent film, successfully answering specific questions about their content. The model was also shown to analyze a 4-hour YouTube video and a 11-hour audio recording, highlighting its potential for complex multimodal understanding.

Gemini 1.5 Pro is built on a new Mixture-of-Experts (MoE) architecture, which Google states makes it more efficient and performant. This architecture allows the model to selectively activate different parts of its neural network for specific tasks, leading to faster processing and reduced computational cost. While the 1 million token context window is available in preview, Google is also offering a smaller 128,000 token context window for general availability, which is still significantly larger than many competing models. The company aims to make the 1 million token window generally available later in 2024.

The release of Gemini 1.5 Pro positions Google at the forefront of AI development, particularly in the area of long-context understanding. This capability has profound implications for fields such as legal research, software development, scientific discovery, and content analysis, where processing large volumes of data is often a bottleneck. The model's multimodal capabilities, meaning it can understand and process different types of data including text, images, audio, and video, further enhance its versatility. Google's ongoing investment in AI research and development, exemplified by the Gemini family of models, underscores the company's commitment to pushing the boundaries of artificial intelligence.

Original source — read the full reporting at the publisher:

Read on The Economist

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next