Interestana
Home/News/OpenAI Introduces GPT-4o with Real-Time Voice and Vision
OpenAI••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

OpenAI Introduces GPT-4o with Real-Time Voice and Vision

OpenAI introduced its latest flagship model, GPT-4o, on May 13, 2024, marking a significant advancement in artificial intelligence by integrating real-time voice and vision processing. This new model, pronounced "GPT-4-omni," is designed to be significantly faster and more capable than its predecessors, offering a more natural and intuitive human-computer interaction experience. GPT-4o processes audio, vision, and text inputs and outputs, enabling it to engage in spoken conversations with users, understand visual information, and generate text responses with unprecedented speed. During a live demonstration, OpenAI showcased GPT-4o's ability to engage in a real-time, back-and-forth spoken conversation, responding to user queries and even detecting the user's emotional state through their tone of voice. The model also demonstrated its capacity to interpret visual input, such as identifying objects in an image, solving math problems presented on a whiteboard, and even translating text in real-time from a live video feed. This multimodal capability means GPT-4o can seamlessly switch between understanding spoken language, interpreting visual cues, and generating appropriate text or spoken responses, all within milliseconds. The company highlighted that GPT-4o achieves GPT-4 Turbo-level intelligence but is 50% faster and significantly more cost-effective, with API pricing reduced by half. This cost reduction is expected to make advanced AI capabilities more accessible to a wider range of developers and businesses. Furthermore, GPT-4o is being rolled out to ChatGPT users, with free users gaining access to features previously exclusive to paid subscribers, including access to GPT-4o's capabilities, albeit with usage limits. The new model's enhanced instruction following and custom voice generation features are also being integrated into the API, allowing developers to build more sophisticated and personalized AI applications. OpenAI stated that the "o" in GPT-4o stands for "omni," signifying its universal capability across different modalities. The development of GPT-4o represents a strategic move by OpenAI to democratize access to cutting-edge AI technology and foster innovation across various sectors, from education and customer service to creative arts and scientific research. The company emphasized its commitment to safety and responsible AI development, noting that GPT-4o has undergone extensive safety testing and will be rolled out gradually to ensure its stable and beneficial integration into society. The enhanced responsiveness and understanding capabilities of GPT-4o are anticipated to revolutionize how people interact with technology, making AI assistants more helpful, engaging, and integrated into daily life.

Original source — read the full reporting at the publisher:

Read on OpenAI

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next