Player not loading? Watch on YouTube
This tutorial walks through FaceFusion setup and a video face swap using a single portrait and a dance clip. The speaker creates and activates a fresh Conda environment in the Anaconda prompt, clones the repository from GitHub and installs the dependencies before launching the browser interface.
For the NVIDIA GPU setup, the tutorial specifies CUDA 12.9 and cuDNN 9.10, with TensorRT as an additional performance option. The speaker also describes OpenVINO as an alternative for an Intel CPU. These are the installation choices presented in the tutorial, rather than a complete hardware compatibility guide.
In the interface, the speaker enables the face swapper and face enhancer, uploads the portrait and target video, and keeps the default models. A preview shows a swapped frame. The reference-face selector lets users choose which face to replace when a video contains multiple faces.
The local AI workflow still needs downloads: FaceFusion fetches models and weights on first launch, and selecting an enhancer or a different model can trigger more downloads. Viewers should account for that before planning offline use. The speaker estimates about 30 minutes for this render, then presents the original and processed clips side by side. That estimate applies to the demonstrated job; the transcript does not establish a general processing speed.