Player not loading? Watch on YouTube
This Russian-language tutorial uses CosyVoice 3 for local speech generation in ComfyUI. The presenter prefers its Russian results among the TTS models he has tried, while limiting that judgment to his own examples.
Setup installs the FL-CosyVoice3 custom nodes and loads a supplied JSON workflow. The model loader downloads the selected weights from Hugging Face or ModelScope and chooses CPU or GPU execution. Reference-based cloning also downloads Whisper to transcribe the sample audio.
The examples cover cloning a voice, mixing languages, replacing a recording's voice and creating dialogue with up to four reference speakers. Saved speaker files let the workflow reuse a voice without supplying its audio each time. Emotion instructions are demonstrated separately.
The walkthrough also shows practical limits. Reference and voice-conversion inputs are restricted to 30 seconds in this workflow. Generated phrases can have audible artifacts at their joins. English or Chinese emotion instructions work differently from Russian instructions, which become spoken text. A line break between a speaker label and its dialogue causes a parsing error. When a word from the supplied reference text leaks into a saved voice's output, the presenter removes that text and lets Whisper transcribe the sample instead.