Player not loading? Watch on YouTube
This tutorial and review tests Seed-VC voice conversion through a custom Google Colab notebook. The speaker describes the open source model as supporting zero-shot conversion with a short target voice sample. The demonstrated setup runs in the cloud on a T4 GPU; it does not show how to run models locally.
The walkthrough opens the notebook from GitHub, selects the T4 runtime, and runs five cells for setup, dependencies, model downloads, the conversion script, and a public Gradio interface. The speaker estimates four to five minutes for startup. In the interface, users upload source audio and a clean target speaker sample, then choose V1 speech, V1 singing, or V2 speech. Vocal separation stays off for spoken audio and turns on for songs, where the notebook separates vocals, converts them, and mixes them back with the instrumentals. The speaker says the singing engine outputs at 44.1 kHz.
The speech comparisons use the same source clip with several target voices. The presenter prefers V2 for the Trump example but finds V1 more natural for Sam Altman, so the results do not establish one engine as consistently better. Singing tests use a roughly 30-second Persian song attributed to Gemini's Lyra model. The presenter finds those conversions less convincing and concludes that complex melodies remain a weakness in these examples.