Beats Suno v5 on SongBench and Runs on a 24GB GPU: m-a-p's YuE2 Makes Open Music Generation Frontier-Grade
The open-source m-a-p collective just shipped YuE2-3B, an open-weight music model that tops WildSongBench with an editable ABC score, agentic editing, cover generation, and 71-second songs on an RTX 4090.
For most of this year, the story of AI music generation has been a story about closed platforms: Suno, Udio, Mureka, and MiniMax’s Music line hoarding the frontier while open-weight challengers clustered a visible distance behind. That gap just closed. The Multimodal Art Projection (m-a-p) collective — the community lab behind the original YuE model — has released YuE2-3B, an open music generation model that doesn’t just approach the proprietary leaders: on the WildSongBench composite benchmark, it posts the highest SongBench average of any evaluated system, open or closed, including Suno v5.
What the numbers actually say
Precision matters here, because the headline invites overselling. On 192 WildSongBench prompts, YuE2 run in best-of-8 selection mode achieves a SongBench average of 6.9632, edging out Suno v5 at 6.8721 and Kunlun’s Mureka 9 at 6.9377. The default two-candidate configuration scores 6.7316 — below the Suno v5 line but well clear of every other open model, including MiniMax Music 3 (6.2830), LeVo 2 (6.3247), and ACE-Step 1.5 (6.0118).
SongBench average aggregates seven quality dimensions, and YuE2’s musicality score (6.27 best-of-8) is the best in the tables. But the per-metric breakdown shows where the open model still concedes ground: Suno v5 retains the lead on MuLan style-text alignment (0.5428 vs 0.5051), on AllMusicCaps caption adherence (0.4353 vs 0.3980), and on phoneme error rate (8.10% vs 9.79%) — meaning Suno still follows your prompt more faithfully and enunciates more cleanly. YuE2 wins on musicality and overall song quality; the proprietary leaders win on instruction-following precision. A benchmark win with best-of-8 selection against single-candidate commercial systems is also not an apples-to-apples comparison, and the team is transparent about the differing protocols.
The second benchmark is arguably more impressive because it isn’t close. On SHS100K zero-shot cover generation — re-rendering an existing song in a new style while preserving its identity — YuE2 with a full score achieves CLEWS mAP of 0.647 and Hit@1 of 71.3%, against 0.419 and 48.4% for SongEcho and near-zero for ACE-Step 1.5. Nothing else in either table does covers this well.
The architectural bet: editable scores instead of vibes
YuE2’s most distinctive design decision is invisible in the benchmark tables but obvious the moment you use it: generation is planned through an editable symbolic score in ABC notation. One AR–NAR Mixture-of-Transformers backbone writes the score and semantic tokens, then produces acoustic latents via flow matching, which a VAE decodes into 48 kHz stereo audio. Three planning modes trade control for autonomy: cot="full" plans both melody and chords, cot="melody" plans melody only (recommended for covers), and cot="off" generates directly from the prompt like a conventional lyric-to-song model.
This turns music generation from a slot machine into an editor. You can export the plan, modify a chord progression or melodic phrase by hand in ABC, and regenerate with the revision — strict reharmonization with the melody preserved, or broader adaptation with lyric and style changes. Because the score is plain text, it’s also a natural interface for agents, and m-a-p leans into that with an agentic editing workflow: you give an agent the score, prompt, and lyrics plus your feedback (“reharmonize the bridge, add a saxophone solo, swap the second verse to English”), and it revises the score, style, and lyrics before YuE2 renders the next version. The project page demonstrates the process on a song called The Last Train, walked through 9 editing steps and 14 versions, migrating from Mandarin pop to English jazz with modern harmony — with the full conversation, scores, and audio at every step inspectable.
The cover-song pipeline is similarly pragmatic: transcribe the source recording with m-a-p’s newly released SheetSage2 to get a melody ABC, gather or transcribe lyrics, pick a target style, and generate. The demos range from Auld Lang Syne as jazz-funk to Jingle Bells as heavy metal.
Practical engineering
The local experience is unusually well-scoped for an open music model. YuE2-3B needs a 24GB NVIDIA GPU with BF16 support and 24GB of host RAM, running unquantized; a 3.6-minute song renders in 71 seconds on an RTX 4090 with peak VRAM around 11.2 GiB — meaning the requirement is really about total memory headroom, and maximum-context testing peaked at 14.08 GiB. For serving, m-a-p publishes a vLLM 0.19 deployment recipe on an H800 that reaches 3,231 LM tokens/s at concurrency 32, translating to roughly 373 songs per hour from a single node. Installation is a wheel from Hugging Face, a YuE2Pipeline.from_pretrained() call, and you’re generating; planning, semantic generation, synthesis, and decoding are also exposed as separate APIs for pipeline builders.
The release is a family, not a single checkpoint: YuE2-3B, two VAE decoders (YuE2-Vae for perceptual quality, YuE2-Vae-legacy for reproducing benchmark numbers), the MERT-v2 music understanding models, the SheetSage2 transcriber, and the WildSongBench dataset itself.
The catch
Two caveats temper the celebration. First, the weights ship under CC BY-NC 4.0 — non-commercial. Just as observers noted when MiniMax open-weighted Music 3 last month, “open-weight” and “usable in a product” are different claims, and anyone hoping to build a Suno competitor on YuE2 will need a separate license that doesn’t exist yet. Second, there’s no YuE2 technical report yet; the model card asks researchers to cite the original YuE paper (arXiv:2503.08638) for now, so the training recipe behind the score-planning approach remains undocumented.
Even so, the trajectory is hard to argue with. Eighteen months ago, open music generation produced curiosities; today it tops a composite benchmark that includes every major commercial system, runs on a gaming GPU, exports its intermediate representation as human-readable sheet music, and treats editing as a first-class, agent-drivable workflow rather than a regeneration lottery. The closed platforms still hold the edge in prompt adherence and vocal clarity — and the commercial questions are real — but the open frontier in music generation has, for the first time, legitimately arrived.