YuE2-3B: Generate and Edit Full Songs on a 24GB GPU
YuE2-3B turns lyrics and a style prompt into a full 48 kHz stereo song. You can edit the melody and chords as an ABC score and render it again.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
YuE2-3B is an open music generation model from the Multimodal Art Projection (m-a-p) team, the group behind the original YuE. You give it lyrics and a style prompt, and it returns a complete song with vocals and accompaniment as 48 kHz stereo audio. What sets it apart is that the model first writes a symbolic plan: an ABC score with the melody and chords. Then it renders audio from that plan. You can export the score, change it by hand or with an LLM agent, and generate the song again. The same mechanism handles covers: transcribe an existing song’s melody, pass it in, and ask for a new style. The weights are released under CC BY-NC 4.0, so you can’t use them commercially.
Key specs
- Output: full songs with vocals and accompaniment, 48 kHz stereo. Lyrics in English and Mandarin.
- Architecture: one AR–NAR Mixture-of-Transformers backbone writes the score and semantic tokens. Flow matching then produces acoustic latents, and a separate VAE (YuE2-Vae) decodes them to audio.
- Planning modes: melody + chords (
cot="full", the default), melody only (cot="melody"), or no plan (cot="off"). - Hardware: Linux, Python 3.10+, a 24GB NVIDIA GPU with BF16 support, and 24GB of available host RAM. No quantization is needed.
Speed on the Hugging Face package, as reported in the card (BF16 AR/NAR, FP32 VAE, default YuE2-Vae, averaged over warm requests):
| GPU | CoT | Generation / audio length | Peak VRAM |
|---|---|---|---|
| RTX 4090 24GB | full | 71.04 / 214.85 s | 11.18 GiB |
| RTX 4090 24GB | melody | 68.68 / 214.67 s | 11.02 GiB |
| RTX 4090 24GB | off | 57.91 / 196.88 s | 11.09 GiB |
Maximum-context testing peaked at 14.08 GiB. The card also lists a separate vLLM 0.19 serving runtime on an H800. It reaches 373.53 songs/hour at an AR concurrency limit of 32, but the card doesn’t include setup instructions for it.
On WildSongBench (192 prompts), YuE2 best-of-8 has the highest SongBench average in the card’s comparison at 6.9632. Suno v5 scores 6.8721, Mureka 9 scores 6.9377, and Suno v6 scores 6.5562. Standard YuE2, which picks from two candidates, scores 6.7316. That beats Suno v6 but trails Suno v5 and Mureka 9. This isn’t a like-for-like comparison: best-of-8 picks from eight candidates and ranks them by Musicality first, while Suno v6 and v6 Wild get two candidates with lower-PER selection, and earlier Suno versions keep their delivered-candidate protocols. On SHS100K zero-shot covers, YuE2 with a full score gets a CLEWS mAP of 0.647, compared with 0.419 for SongEcho and 0.024 for ACE-Step 1.5.
Install
The card ships a custom inference wheel inside the model repo. It pins huggingface-hub:
python -m pip install huggingface-hub==0.36.2hf download m-a-p/YuE2-3B yue2_infer-0.1.5-py3-none-any.whl --local-dir .python -m pip install ./yue2_infer-0.1.5-py3-none-any.whlRun it
This example uses the style, lyrics and seed from the card’s Mandarin funk / nu-disco demo:
import jsonfrom pathlib import Pathfrom huggingface_hub import hf_hub_downloadfrom yue2 import YuE2Pipeline
repo = "m-a-p/YuE2-3B"pipe = YuE2Pipeline.from_pretrained(repo, device="cuda")
prompt_path = hf_hub_download(repo, "examples/tonight-awake.json")demo = json.loads(Path(prompt_path).read_text(encoding="utf-8"))style, lyrics = demo["style"], demo["lyrics"]
song = pipe(style=style, lyrics=lyrics, cot="full", seed=demo["seed"])song.save("song.flac")song.save_artifacts("outputs/song") # ABC, tokens, latents, audio and settings
pipe.close()save_artifacts writes the generated score to outputs/song/score.abc, which is the starting point for editing. The card doesn’t document the lyric format in prose. Open examples/tonight-awake.json to see how the authors structure sections, and copy that layout.
The card lists cfg_scale=1.2 as an option for stronger text guidance. Semantic CFG defaults to 1.0 for full/melody and 1.01 for off.
The card also mentions a command-line tool, yue2 generate and yue2 batch. It shows only the --quiet flag for these and doesn’t document the other options.
Cover an existing song
For covers, the authors recommend melody-only mode. The workflow has three steps:
- Transcribe the original recording with SheetSage2. Save the melody ABC without chord symbols as
melody.abc. - Get the lyrics, either from an online source or by transcribing them with an ASR model such as Qwen3-ASR. Organize them into sections that match the recording and save them as
cover_lyrics.txt. - Run YuE2 with the target style.
The card says to run the transcription tools in their own environments. The YuE2 step looks like this:
from pathlib import Pathfrom yue2 import YuE2Pipeline
pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda")
cover = pipe( style="Jazz-funk, warm lead vocal, Rhodes piano, electric bass, tight drums", lyrics=Path("cover_lyrics.txt").read_text(encoding="utf-8"), abc=Path("melody.abc").read_text(encoding="utf-8"), cot="melody", seed=831001,)cover.save("cover.flac")cover.save_artifacts("outputs/cover")
pipe.close()To keep the original or an edited harmony, leave the chord symbols in the ABC and use cot="full".
The SHS100K numbers show the tradeoff between staying close to the source and sounding good:
- Full score: best song identity (CLEWS Hit@1 71.3%).
- Score without chords: slightly lower identity (67.3%) but higher Musicality (5.490 vs 5.104).
- No score: essentially loses the source song (0.3% Hit@1).
Edit the score, then re-render
This is the main reason to pick YuE2 over a prompt-only model. You edit the ABC score yourself or have an LLM agent revise it, and then render it again, optionally with a new style.
The documented way to get an editable score is to run create.py first, which generates the full song, and then copy outputs/song/score.abc to edited.abc. The card also shows pipe.plan(), which produces a plan before any audio is generated, and plan.save(). But it doesn’t say which files plan.save() writes, so this post doesn’t build on it.
The card’s example keeps the lyrics and seed from the first run but changes the style prompt too, so the re-render reflects both your score edits and the new style. The card doesn’t say whether keeping the seed gives you a controlled comparison after you change the ABC.
import jsonfrom pathlib import Pathfrom huggingface_hub import hf_hub_downloadfrom yue2 import YuE2Pipeline
# Before running: run create.py, copy outputs/song/score.abc to edited.abc,# and edit edited.abc.edited_path = Path("edited.abc")if not edited_path.exists(): raise SystemExit( "edited.abc not found. Run create.py, copy outputs/song/score.abc " "to edited.abc, edit it, then run this script again." )
repo = "m-a-p/YuE2-3B"pipe = YuE2Pipeline.from_pretrained(repo, device="cuda")
prompt_path = hf_hub_download(repo, "examples/tonight-awake.json")demo = json.loads(Path(prompt_path).read_text(encoding="utf-8"))lyrics = demo["lyrics"]
edited_style = ( "Jazz, expressive lead vocal, piano, tenor saxophone, upright bass, " "brushed drums, no guitar, spacious modern harmony")song = pipe( style=edited_style, lyrics=lyrics, cot="full", seed=demo["seed"], abc=edited_path.read_text(encoding="utf-8"),)song.save("edited.flac")song.save_artifacts("outputs/edited")
pipe.close()When you hand the score to an agent for strict reharmonization, the card recommends two instructions. Tell the agent to keep the melody pitches and rhythm unchanged, and to check sustained notes against the new chords. For looser rewrites, allow specific melody or lyric changes. The project’s demo page shows a 9-step, 14-version editing session that goes from Mandarin pop to English jazz.
To control the stages yourself, the pipeline exposes pipe.plan() → pipe.generate_semantic(plan) → pipe.synthesize(semantic) → pipe.decode(latents).
Generate several takes
The benchmark results depend on candidate selection: standard YuE2 picks from two candidates and best-of-8 picks from eight. Each pipeline call produces one candidate, and the HF package generates one song at a time. If you want that quality, generate several seeds and choose the best take yourself.
The style and lyrics.txt below are placeholders. Format your lyrics file with the same section layout as examples/tonight-awake.json.
from pathlib import Pathfrom yue2 import YuE2Pipeline
pipe = YuE2Pipeline.from_pretrained("m-a-p/YuE2-3B", device="cuda")
style = "English rock, gritty male lead vocal, distorted guitars, bass, live drums"lyrics = Path("lyrics.txt").read_text(encoding="utf-8")
out = Path("takes")out.mkdir(exist_ok=True)
for seed in [101, 202, 303, 404]: song = pipe(style=style, lyrics=lyrics, cot="full", seed=seed) song.save(str(out / f"take_{seed}.flac")) song.save_artifacts(str(out / f"take_{seed}"))
pipe.close()On an RTX 4090, a warm full-CoT take of about 3.6 minutes averages roughly 71 seconds. That timing leaves out initial path resolution and saving, so each real call takes somewhat longer. It’s also an average over warm requests, so it doesn’t cover a cold first call, and the card gives no cold-start time. For selection, the card’s standard setting picks the candidate with the lower phoneme error rate out of two. Best-of-8 ranks by Musicality, then prompt adherence, then phoneme error rate. The card says selection is a separate step from the pipeline call and doesn’t say whether the scoring tools come with the package. Plan to listen to the takes yourself or bring your own scorers.
Gotchas
- Non-commercial license. The weights are CC BY-NC 4.0, so you can’t use the model in a paid product. Third-party code has its own licenses, listed in
THIRD_PARTY_NOTICES.md. - Platform requirements. The card targets Linux and an NVIDIA GPU with BF16 support. Plan for 24GB of VRAM and 24GB of free host RAM, even though measured peaks are 11–14 GiB.
- Pinned dependency. Install
huggingface-hub==0.36.2as shown. The package isn’t a standard Transformers pipeline; you load it through theyue2wheel. - Melody mode keeps chords.
cot="melody"doesn’t remove chord symbols from your ABC. Strip them yourself for covers. - Two VAEs. The card says the default
YuE2-Vaedelivers better perceptual audio quality.YuE2-Vae-legacyscores higher on benchmark musicality and was used for the WildSongBench and SHS100K results. The speed and VRAM numbers use the default YuE2-Vae. Passvae="m-a-p/YuE2-Vae-legacy"tofrom_pretrainedif you need to reproduce the benchmarks. - Lyric accuracy isn’t the best. YuE2’s phoneme error rate (8.44%) is higher than MiniMax Music 3 (6.27%), ACE-Step 1.5 (7.46%), and every Suno version in the card (5.80–8.10%). Best-of-8 is higher still (9.79%) because it ranks candidates by Musicality before PER. PER is measured with ASR, so read it as a rough signal of lyric errors.
- One song at a time. The HF package renders songs one after another. The card mentions a
yue2 batchcommand but doesn’t document how it works. The vLLM serving setup, which handles concurrent requests, is separate, and the card doesn’t explain how to deploy it. - Clean up and quiet logs. Call
pipe.close()when you’re done. The card showsYuE2Pipeline.from_pretrained(repo, progress=False)to turn off progress messages, or--quieton the CLI.
When to pick it
Pick YuE2-3B when you want full-length songs with vocals from open weights on a single 24GB GPU. It’s especially strong if you want control over the music itself and not only the prompt. The editable ABC score makes it a good fit for agent-driven revision loops, reharmonization experiments, and style covers that keep a song recognizable. On the card’s benchmarks it beats every other open model listed on SongBench average. Its best-of-8 setting scores above the Suno versions the card tested, but under a different candidate-selection protocol, so it isn’t a like-for-like win.
Don’t pick it for anything commercial: the license rules that out. It’s also not a good choice if accurate lyrics are your top priority, if you need instrumental-only or short sound-effect generation (the card doesn’t cover those), or if you need high-throughput serving without building the vLLM setup yourself. The benchmark results also depend on picking the best of 2 or 8 candidates, so plan to generate several takes per song.
Related
Run Kokoro-82M: Apache-Licensed Text-to-Speech in a Few Lines
Kokoro-82M is an 82M-parameter open-weight TTS model. How to install it, generate speech, and use it for narration, pronunciation fixes, and captions.
hexgrad/Kokoro-82M
Run Audio8 ASR Infinite for 24/7 Chinese and English Transcription
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite
AuK: Clone, Edit, and Clean Up Speech Through One Instruction Interface
Tencent's MIT-licensed AuK does zero-shot TTS, speech editing, enhancement, and separation from natural-language instructions.
tencent/AuK