Player not loading? Watch on YouTube
This tutorial shows how to run YuE2 in ComfyUI for local AI music generation. The speaker demonstrates songs in several languages, covers based on reference melodies, and iterative edits through an AI agent. YuE2 first produces editable ABC notation for vocals and instruments, then uses that score to generate audio.
Setup starts with updating ComfyUI and importing the supplied JSON workflow. The walkthrough places a 1.4 GB audio encoder in models/audio_encoders and a checkpoint in models/checkpoints. The speaker presents a 7.8 GB BF16 checkpoint as suitable for roughly 8 GB of VRAM, and recommends the 3.96 GB quantized option for less than 4 GB. The suggested 12 GB threshold for regular versus tiled decoding is explicitly the speaker's estimate.
The workflow accepts style prompts and lyrics with section tags. It also exposes duration, seed, steps and CFG controls. One example runs longer than intended, prompting advice to reduce the maximum duration. For instrumentals, the speaker enters bracketed instrument directions in the lyrics field. Reference audio can supply either a melody or a full-song reference, with new styles and lyrics.
The speaker cites benchmark comparisons with Suno, but these do not establish universal quality rankings. Licensing also needs attention: the tutorial identifies Apache 2 licensing for code and documentation, while model weights carry a separate Creative Commons non-commercial license.