Interestana
Home/News/MiniMax H3 Generates 15-Second 2K Video With Stereo Audio
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

MiniMax H3 Generates 15-Second 2K Video With Stereo Audio

MiniMax released MiniMax H3 on July 31, 2026, a general-purpose omni-modal generation model designed to process and generate video with native stereo sound. Unlike previous text-to-video models that often relied on separate expert models for different functionalities, MiniMax H3 integrates text, images, video, and audio into a unified context for generation. This approach allows the model to understand and execute complex prompts that reference elements from multiple modalities, such as camera movements from one video, character actions from an image, and vocalizations from an audio file, all expressed through natural language.

The core specifications for MiniMax H3 include the ability to generate video output at 2K resolution, with clip durations ranging from 4 to 15 seconds, exclusively in integer durations. The model's architecture consolidates various video generation tasks, including text-to-video, image-to-video, first-and-last-frame generation, subject reference, motion reference, and video editing, into a single pretraining paradigm. This unified framework enables the model to express reference and editing relationships naturally within prompts.

MiniMax H3 is currently available via an API under the model ID MiniMax-H3 and within the consumer Hailuo AI application. The company positions this model for a wide array of industries, including advertising, branding, e-commerce, product design, UI/UX, and gaming, as well as for film pre-visualization and retail catalog media. Potential applications span ad variant generation, product and listing videos, animated posters, film title sequences, website hero loops, character-consistent game cinematics, and video-to-video motion transfer.

The API for MiniMax H3 offers three primary entry modes: text-to-video, first/last-frame image-to-video, and reference generation. The generation process is asynchronous, involving a three-step flow: creating a task, polling for the task ID, and then downloading the generated content from a provided URL. Input limitations for the API include a maximum of 9 reference images, up to 3 reference video clips totaling no more than 15 seconds, and up to 3 reference audio clips. Audio inputs require an accompanying image or video. The total number of mixed inputs is capped at 12 files, with a maximum prompt length of 7,000 characters and a request body size limit of 64 MB.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next