Interestana
Home/News/MiniMax Releases Open-Weights Music Model Generating Five-Minute Songs
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

MiniMax Releases Open-Weights Music Model Generating Five-Minute Songs

MiniMax released MiniMax-Music3 on an unspecified date, an open-weights text-to-music model designed to generate complete songs from lyrical and descriptive inputs. This model accepts two distinct inputs: lyrics that include section tags and a detailed music description. It then generates a full song, up to five minutes in length, in a single pass. The output format is a 32 kHz, 16-bit stereo WAV file. The underlying architecture of MiniMax-Music3 pairs a Hybrid-LM, which comprises an 8 billion parameter Global LLM and a 0.6 billion parameter Local LLM, with a continuous synthesis stack. This synthesis stack is built upon flow matching and a Flow-Variational Autoencoder (Flow-VAE). MiniMax simultaneously released the model's weights, inference code, and three documented serving paths, indicating immediate deployability rather than a research preview. The MiniMax-Music3 Community License permits commercial use, with the stipulation that 'MiniMax-Music3' must be prominently displayed in the product's user interface. Organizations generating over $20 million in aggregate yearly revenue from products utilizing the model must obtain separate written authorization from MiniMax. Furthermore, any entity hosting third-party generation services must implement and maintain safeguards to prevent infringing outputs. The model is positioned for use across various industries, including game development, advertising and brand agencies, short-form video platforms, creator tools, e-learning, podcasting, fitness and wellness applications, retail in-store audio systems, and music-technology Software as a Service (SaaS) providers. Potential applications include generating background scores for user-generated content videos, creating adaptive music for games, producing localized advertising jingles and sonic branding, composing scratch and demo tracks for musicians, enabling mood-conditioned playlist generation, and facilitating offline batch generation for scenarios where per-song API costs are a limiting factor. Architecturally, MiniMax-Music3 integrates a hierarchical autoregressive stack with a continuous synthesis pathway. The training tokenizer employs eight layers of residual vector quantization (RVQ). The topmost layer, a semantic codebook, contains 16,384 entries and is responsible for capturing core musical semantics and structure. The subsequent seven acoustic codebooks each contain 1,024 entries, contributing to the detailed acoustic generation of the music.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next