Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 8 min read

YuE2-3B: Generate and Edit Full Songs on a 24GB GPU

YuE2-3B turns lyrics and a style prompt into a full 48 kHz stereo song. You can edit the melody and chords as an ABC score and render it again.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

YuE2-3B is an open music generation model from the Multimodal Art Projection (m-a-p) team, the group behind the original YuE. You give it lyrics and a style prompt, and it returns a complete song with vocals and accompaniment as 48 kHz stereo audio. What sets it apart is that the model first writes a symbolic plan: an ABC score with the melody and chords. Then it renders audio from that plan. You can export the score, change it by hand or with an LLM agent, and generate the song again. The same mechanism handles covers: transcribe an existing song’s melody, pass it in, and ask for a new style. The weights are released under CC BY-NC 4.0, so you can’t use them commercially.

Key specs

  • Output: full songs with vocals and accompaniment, 48 kHz stereo. Lyrics in English and Mandarin.
  • Architecture: one AR–NAR Mixture-of-Transformers backbone writes the score and semantic tokens. Flow matching then produces acoustic latents, and a separate VAE (YuE2-Vae) decodes them to audio.
  • Planning modes: melody + chords (cot="full", the default), melody only (cot="melody"), or no plan (cot="off").
  • Hardware: Linux, Python 3.10+, a 24GB NVIDIA GPU with BF16 support, and 24GB of available host RAM. No quantization is needed.

Speed on the Hugging Face package, as reported in the card (BF16 AR/NAR, FP32 VAE, default YuE2-Vae, averaged over warm requests):

GPU CoT Generation / audio length Peak VRAM
RTX 4090 24GB full 71.04 / 214.85 s 11.18 GiB
RTX 4090 24GB melody 68.68 / 214.67 s 11.02 GiB
RTX 4090 24GB off 57.91 / 196.88 s 11.09 GiB

Maximum-context testing peaked at 14.08 GiB. The card also lists a separate vLLM 0.19 serving runtime on an H800. It reaches 373.53 songs/hour at an AR concurrency limit of 32, but the card doesn’t include setup instructions for it.

On WildSongBench (192 prompts), YuE2 best-of-8 has the highest SongBench average in the card’s comparison at 6.9632. Suno v5 scores 6.8721, Mureka 9 scores 6.9377, and Suno v6 scores 6.5562. Standard YuE2, which picks from two candidates, scores 6.7316. That beats Suno v6 but trails Suno v5 and Mureka 9. This isn’t a like-for-like comparison: best-of-8 picks from eight candidates and ranks them by Musicality first, while Suno v6 and v6 Wild get two candidates with lower-PER selection, and earlier Suno versions keep their delivered-candidate protocols. On SHS100K zero-shot covers, YuE2 with a full score gets a CLEWS mAP of 0.647, compared with 0.419 for SongEcho and 0.024 for ACE-Step 1.5.

Install

The card ships a custom inference wheel inside the model repo. It pins huggingface-hub:

Terminal window
python -m pip install huggingface-hub==0.36.2
hf download m-a-p/YuE2-3B yue2_infer-0.1.5-py3-none-any.whl --local-dir .
python -m pip install ./yue2_infer-0.1.5-py3-none-any.whl

Run it

This example uses the style, lyrics and seed from the card’s Mandarin funk / nu-disco demo:

create.py
import json
from pathlib import Path
from huggingface_hub import hf_hub_download
from yue2 import YuE2Pipeline
repo = "m-a-p/YuE2-3B"
pipe = YuE2Pipeline.from_pretrained(repo, device="cuda")
prompt_path = hf_hub_download(repo, "examples/tonight-awake.json")
demo = json.loads(Path(prompt_path).read_text(encoding="utf-8"))
style, lyrics = demo["style"], demo["lyrics"]
song = pipe(style=style, lyrics=lyrics, cot="full", seed=demo["seed"])
song.save("song.flac")
song.save_artifacts("outputs/song") # ABC, tokens, latents, audio and settings
pipe.close()

save_artifacts writes the generated score to outputs/song/score.abc, which is the starting point for editing. The card doesn’t document the lyric format in prose. Open examples/tonight-awake.json to see how the authors structure sections, and copy that layout.

The card lists cfg_scale=1.2 as an option for stronger text guidance. Semantic CFG defaults to 1.0 for full/melody and 1.01 for off.

The card also mentions a command-line tool, yue2 generate and yue2 batch. It shows only the --quiet flag for these and doesn’t document the other options.

Cover an existing song

For covers, the authors recommend melody-only mode. The workflow has three steps:

  1. Transcribe the original recording with SheetSage2. Save the melody ABC without chord symbols as melody.abc.
  2. Get the lyrics, either from an online source or by transcribing them with an ASR model such as Qwen3-ASR. Organize them into sections that match the recording and save them as cover_lyrics.txt.
  3. Run YuE2 with the target style.

The card says to run the transcription tools in their own environments. The YuE2 step looks like this:

cover.py
from pathlib import Path
from yue2 import YuE2Pipeline
pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda")
cover = pipe(
style="Jazz-funk, warm lead vocal, Rhodes piano, electric bass, tight drums",
lyrics=Path("cover_lyrics.txt").read_text(encoding="utf-8"),
abc=Path("melody.abc").read_text(encoding="utf-8"),
cot="melody",
seed=831001,
)
cover.save("cover.flac")
cover.save_artifacts("outputs/cover")
pipe.close()

To keep the original or an edited harmony, leave the chord symbols in the ABC and use cot="full".

The SHS100K numbers show the tradeoff between staying close to the source and sounding good:

  • Full score: best song identity (CLEWS Hit@1 71.3%).
  • Score without chords: slightly lower identity (67.3%) but higher Musicality (5.490 vs 5.104).
  • No score: essentially loses the source song (0.3% Hit@1).

Edit the score, then re-render

This is the main reason to pick YuE2 over a prompt-only model. You edit the ABC score yourself or have an LLM agent revise it, and then render it again, optionally with a new style.

The documented way to get an editable score is to run create.py first, which generates the full song, and then copy outputs/song/score.abc to edited.abc. The card also shows pipe.plan(), which produces a plan before any audio is generated, and plan.save(). But it doesn’t say which files plan.save() writes, so this post doesn’t build on it.

The card’s example keeps the lyrics and seed from the first run but changes the style prompt too, so the re-render reflects both your score edits and the new style. The card doesn’t say whether keeping the seed gives you a controlled comparison after you change the ABC.

plan_and_edit.py
import json
from pathlib import Path
from huggingface_hub import hf_hub_download
from yue2 import YuE2Pipeline
# Before running: run create.py, copy outputs/song/score.abc to edited.abc,
# and edit edited.abc.
edited_path = Path("edited.abc")
if not edited_path.exists():
raise SystemExit(
"edited.abc not found. Run create.py, copy outputs/song/score.abc "
"to edited.abc, edit it, then run this script again."
)
repo = "m-a-p/YuE2-3B"
pipe = YuE2Pipeline.from_pretrained(repo, device="cuda")
prompt_path = hf_hub_download(repo, "examples/tonight-awake.json")
demo = json.loads(Path(prompt_path).read_text(encoding="utf-8"))
lyrics = demo["lyrics"]
edited_style = (
"Jazz, expressive lead vocal, piano, tenor saxophone, upright bass, "
"brushed drums, no guitar, spacious modern harmony"
)
song = pipe(
style=edited_style,
lyrics=lyrics,
cot="full",
seed=demo["seed"],
abc=edited_path.read_text(encoding="utf-8"),
)
song.save("edited.flac")
song.save_artifacts("outputs/edited")
pipe.close()

When you hand the score to an agent for strict reharmonization, the card recommends two instructions. Tell the agent to keep the melody pitches and rhythm unchanged, and to check sustained notes against the new chords. For looser rewrites, allow specific melody or lyric changes. The project’s demo page shows a 9-step, 14-version editing session that goes from Mandarin pop to English jazz.

To control the stages yourself, the pipeline exposes pipe.plan() → pipe.generate_semantic(plan) → pipe.synthesize(semantic) → pipe.decode(latents).

Generate several takes

The benchmark results depend on candidate selection: standard YuE2 picks from two candidates and best-of-8 picks from eight. Each pipeline call produces one candidate, and the HF package generates one song at a time. If you want that quality, generate several seeds and choose the best take yourself.

The style and lyrics.txt below are placeholders. Format your lyrics file with the same section layout as examples/tonight-awake.json.

takes.py
from pathlib import Path
from yue2 import YuE2Pipeline
pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda")
style = "English rock, gritty male lead vocal, distorted guitars, bass, live drums"
lyrics = Path("lyrics.txt").read_text(encoding="utf-8")
out = Path("takes")
out.mkdir(exist_ok=True)
for seed in [101, 202, 303, 404]:
song = pipe(style=style, lyrics=lyrics, cot="full", seed=seed)
song.save(str(out / f"take_{seed}.flac"))
song.save_artifacts(str(out / f"take_{seed}"))
pipe.close()

On an RTX 4090, a warm full-CoT take of about 3.6 minutes averages roughly 71 seconds. That timing leaves out initial path resolution and saving, so each real call takes somewhat longer. It’s also an average over warm requests, so it doesn’t cover a cold first call, and the card gives no cold-start time. For selection, the card’s standard setting picks the candidate with the lower phoneme error rate out of two. Best-of-8 ranks by Musicality, then prompt adherence, then phoneme error rate. The card says selection is a separate step from the pipeline call and doesn’t say whether the scoring tools come with the package. Plan to listen to the takes yourself or bring your own scorers.

Gotchas

  • Non-commercial license. The weights are CC BY-NC 4.0, so you can’t use the model in a paid product. Third-party code has its own licenses, listed in THIRD_PARTY_NOTICES.md.
  • Platform requirements. The card targets Linux and an NVIDIA GPU with BF16 support. Plan for 24GB of VRAM and 24GB of free host RAM, even though measured peaks are 11–14 GiB.
  • Pinned dependency. Install huggingface-hub==0.36.2 as shown. The package isn’t a standard Transformers pipeline; you load it through the yue2 wheel.
  • Melody mode keeps chords. cot="melody" doesn’t remove chord symbols from your ABC. Strip them yourself for covers.
  • Two VAEs. The card says the default YuE2-Vae delivers better perceptual audio quality. YuE2-Vae-legacy scores higher on benchmark musicality and was used for the WildSongBench and SHS100K results. The speed and VRAM numbers use the default YuE2-Vae. Pass vae="m-a-p/YuE2-Vae-legacy" to from_pretrained if you need to reproduce the benchmarks.
  • Lyric accuracy isn’t the best. YuE2’s phoneme error rate (8.44%) is higher than MiniMax Music 3 (6.27%), ACE-Step 1.5 (7.46%), and every Suno version in the card (5.80–8.10%). Best-of-8 is higher still (9.79%) because it ranks candidates by Musicality before PER. PER is measured with ASR, so read it as a rough signal of lyric errors.
  • One song at a time. The HF package renders songs one after another. The card mentions a yue2 batch command but doesn’t document how it works. The vLLM serving setup, which handles concurrent requests, is separate, and the card doesn’t explain how to deploy it.
  • Clean up and quiet logs. Call pipe.close() when you’re done. The card shows YuE2Pipeline.from_pretrained(repo, progress=False) to turn off progress messages, or --quiet on the CLI.

When to pick it

Pick YuE2-3B when you want full-length songs with vocals from open weights on a single 24GB GPU. It’s especially strong if you want control over the music itself and not only the prompt. The editable ABC score makes it a good fit for agent-driven revision loops, reharmonization experiments, and style covers that keep a song recognizable. On the card’s benchmarks it beats every other open model listed on SongBench average. Its best-of-8 setting scores above the Suno versions the card tested, but under a different candidate-selection protocol, so it isn’t a like-for-like win.

Don’t pick it for anything commercial: the license rules that out. It’s also not a good choice if accurate lyrics are your top priority, if you need instrumental-only or short sound-effect generation (the card doesn’t cover those), or if you need high-throughput serving without building the vLLM setup yourself. The benchmark results also depend on picking the best of 2 or 8 candidates, so plan to generate several takes per song.

Related