Player not loading? Watch on YouTube
The video introduces three local text-to-speech nodes for ComfyUI: XTTS v2, Kokoro and MeloTTS. Each node passes text to an external Python script, which generates an audio file and returns it to the workflow. The presenter describes this as a foundation for local AI voice workflows without external APIs or cloud services.
The setup keeps each TTS system in a dedicated Python virtual environment, separate from ComfyUI. The presenter explains that differing dependencies, Torch versions and audio libraries can otherwise cause conflicts. The node launches its script with that environment's Python interpreter. Its controls include the text, script path, Python environment and output folder.
The comparison reflects the presenter's preferences rather than measured benchmarks. They favor XTTS for voice cloning, reference audio and consistent narrators. They describe Kokoro as lightweight and fast for quick tests and automation, and MeloTTS as easy to integrate with multilingual support.
Examples cover narration and talking characters. The presenter also previews a future workflow that uses an LLM to plan scenes, prepares narration and visual prompts, tracks outputs, and assembles a longer video. The supplied description links the three node repositories, and the presenter says they include example workflows.