MiniMax Open-Sources Music 3: Complete Five-Minute Songs From One Prompt
MiniMax Music 3 pairs an 8B Global LLM with a 0.6B Local LLM and a flow-matching synth stage to generate full five-minute songs — and the weights are on Hugging Face.
For years, AI music generation has had a hard ceiling: most tools could produce a catchy 60-to-90-second clip, but full-length tracks fell apart — vocals drifted, arrangements lost the plot, and by the second chorus the model had forgotten the melody it started with. MiniMax’s new Music 3 model, released with open weights on August 13, 2026, is a direct attempt to break that ceiling. Given a creative concept and optional lyrics, it composes, arranges, performs, and produces a complete song of up to five minutes in a single generation pass — and anyone can download the weights from Hugging Face and run it themselves.
The release matters on two fronts at once. Technically, it is one of the most interesting generative-audio architectures shipped this year, replacing the pure discrete-token decoding of earlier systems with a continuous hidden-state synthesis path. Strategically, it extends a now-familiar pattern: another frontier-grade creative model coming out of a Chinese lab with open weights, while American competitors keep their music and video generators behind closed, paid APIs. If 2026 is the year open-weights models caught up with closed ones, Music 3 is one of the clearest data points in media generation.
What Music 3 actually does
Music 3 is a text-to-music model with an unusual degree of control. It takes two inputs:
- Lyrics, including explicit section tags such as
[Verse],[Pre-Chorus],[Chorus],[Bridge],[Instrumental],[Solo], and[Outro]— so you decide the song’s structure, not the model. - A music description, and for precise control a Structured Caption with three sections: Global Metadata (genre, BPM, key, scale, emotional progression, production profile), Vocal Details (gender, timbre, performance style, harmony, backing vocals, effects), and Arrangement (instruments, section-level instrument evolution, groove, bass, percussion, textures, spatial effects).
The output is 32 kHz, 16-bit stereo WAV audio — a finished song, not a loop to be stitched together. Crucially, the model holds musical identity together across the full track: themes, rhythm, vocal identity, and arrangement progression persist from intro to outro, which is exactly where short-clip generators historically collapsed.
MiniMax also shipped a template-based Prompt Enhancement System — including an official music-caption-rewriter skill installable via npx skills add MiniMax-AI/MiniMax-Music3 — that expands a plain-language description (“a warm acoustic pop song with intimate female vocals”) into a full three-part structured caption using proper musical terminology. The intent is clear: let hobbyists exercise producer-grade control without learning producer vocabulary.
Inside the architecture
The technical design is the most compelling part of the release, and it explains why the long-form coherence works. Music 3 has three components: a tokenizer, a hybrid language-model stack, and a continuous synthesis stage.
The tokenizer uses eight layers of Residual Vector Quantization (RVQ), with a deliberate division of labor. The first, semantic codebook holds 16,384 entries and captures core musical structure; the remaining seven acoustic codebooks (1,024 entries each) encode progressively finer acoustic residuals. Training is staged — the semantic codebook is optimized first to build a stable backbone, then all eight codebooks train jointly. This separation keeps long-sequence prediction stable instead of overloading one token stream with both structural and fidelity information.
The Hybrid-LM splits the modeling problem in two. An 8B Global LLM (initialized from Qwen3-8B; MiniMax’s research post says Qwen3.5-8B, so the exact base checkpoint is worth treating as unsettled) predicts the semantic codebook frame by frame while maintaining full-song context — the model’s sense of where the song is and where it is going. A randomly initialized 0.6B Local LLM predicts the acoustic codebooks within each frame along the depth axis, restoring fine-grained sound detail. Training again proceeds in stages: global alignment first, then full-parameter joint training of both models.
The synthesis stack is where Music 3 departs from convention. Instead of decoding audio directly from discrete tokens — the standard approach that loses information at quantization — the model fuses the continuous hidden states of both LLMs and feeds them through a 2.4B flow-matching module, with a 123M Flow-VAE decoder (adapted from MiniMax’s speech stack and retrained for music’s dynamic range) producing the final waveform. The full path is: fused LLM features → flow matching → VAE latent → Flow-VAE decoder → 32 kHz stereo audio. Continuous representations carry more acoustic information than discrete tokens can, and it shows up precisely where earlier models failed: vocal articulation, instrumental texture, and temporal continuity across minutes rather than seconds.
Notably, the discrete tokenizer is a training-time tool only — at inference, waveform synthesis runs entirely off the fused hidden states.
Running it yourself
Three documented paths exist, covering everyone from hobbyists to production deployments:
- SGLang-Omni is the reference serving stack (
sgl-omni serve --model-path MiniMaxAI/MiniMax-Music3), exposing the shared speech API. The GitHub repo specifies a two-GPU setup: GPU 0 handles Qwen3 and RVQ autoregressive generation, GPU 1 handles flow matching and VAE decoding. Generation exposesmax_new_tokensat 25 audio frames per second, ending early on an end-of-audio token. - Diffusers offers a modular pipeline that fits under 24 GB VRAM at full precision, roughly 22 GB with automatic CPU offload, and as low as 8 GB with leaf-level group offloading — meaning a single consumer GPU can run a frontier music model.
- ComfyUI shipped a native text-to-music template on day one, with repacked FP16/INT8 weights from Comfy-Org, which is how the model reached the open-source creative community within hours of release.
The licensing deserves attention. The MiniMax-Music3 Community License permits commercial use, but requires displaying “MiniMax-Music3” prominently in the product UI, and organizations whose aggregate yearly revenue from products built on it exceeds US$20 million must obtain a separate commercial license. It is open, but not OSI-open — closer in spirit to Meta’s Llama Community License than to Apache 2.0.
Why this matters
The competitive context makes the release sharper. Western AI music tools — Suno, Udio, and their peers — remain closed, subscription products, and their catalogs of licensing lawsuits with record labels are still working through courts. Suno, notably, is capping Pro downloads at 20 songs per month as of September 3, which makes the arrival of a genuinely competitive open-weights alternative unusually well timed. ARIA, the Australian recording industry body, moved this week to ban fully-AI music from its charts — a reminder that institutional pushback against synthetic music is rising at the same moment the technology gets dramatically more capable and more accessible.
Music 3 also fits MiniMax’s own trajectory. The company has spent 2026 shipping open media models at a striking pace — the H3 video model in July, Music 3 in August — building out a full-stack creative portfolio where the weights themselves are the marketing. For studios, game developers, and independent creators with privacy constraints or high volume needs, self-hostable music generation at this quality changes the economics of scoring, soundtracks, and jingles. The 8 GB offload floor puts it within reach of a gaming rig.
The open question is provenance and training data. As with every generative music model, what the model learned from — and whether rightsholders come after open weights the way they went after closed services — remains unresolved. But as an engineering artifact, Music 3 marks the moment open-weights music generation stopped being a demo and became production tooling. Five minutes, one pass, your lyrics, your structure — and the model that sings it lives on your own hardware.
Sources
- [1] https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model
- [2] https://huggingface.co/MiniMaxAI/MiniMax-Music3
- [3] https://github.com/minimax-ai/minimax-music3
- [4] https://blog.comfy.org/p/minimax-music-3-state-of-the-art
- [5] https://www.marktechpost.com/2026/08/17/minimax-releases-minimax-music3/