Player not loading? Watch on YouTube
This review tests Zonos 2 text-to-speech and voice cloning through Saganaki22's ComfyUI custom-node workflow and converted model files. The speaker describes Zyphra's model as open source under Apache 2.0, with 900 million active parameters and 8 billion total parameters in a sparse mixture of experts architecture.
The tests compare FP8 and BF16, described as roughly 10 GB and 15 GB downloads. The speaker hears no clear quality advantage from BF16 in the cloning examples. Basic speech generation produces different voices between runs, and the speaker finds it unreliable. One sample cuts off at a max new tokens setting of 1024; increasing it to 2048 allows the passage to finish.
Voice cloning gets mixed results. Tests with temperatures of 0.7 and 0.4 still leave differences in accent and vocal character. The speaker is more positive about multilingual generation, testing French, Japanese and Kurdish and reporting that the model retains the reference voice while adapting pronunciation to the language.
The setup section covers installing missing ComfyUI nodes and placing model files and supporting DAC and speaker encoder files in the expected folders. Automatic model downloads failed in the speaker's setup, so the tutorial uses manual downloads. The results show usable multilingual examples alongside inconsistent basic voices and imperfect accent matching.