Player not loading? Watch on YouTube
This tutorial demonstrates Bernini, a video editing model that the presenter describes as open source and released by a ByteDance team. Using Wan2GP on an RTX 5090 hosted on RunPod, the presenter edits an existing clip through text prompts. The implementation shown omits the planner component described in the model's paper.
The setup selects Wan 2.2, then Bernini R, and applies the December 17, 2025 lightning text-video preset with two phases and four steps. The presenter disables reference images, uploads a video and prompts the model to replace a basketball with a ball of fire. Output is set to 720p at 16:9, with CFG set to 1 by the preset. In these tests, 1080p produced worse results, so the presenter prefers 720p followed by upscaling after reviewing the output.
The five-second fireball edit took about four minutes on the demonstrated hardware. The presenter reports that much of the original scene stayed intact, though the basketball rim appeared faded. No mask is used; the presenter says masking had no apparent effect in their tests.
A hamburger reference-image edit gives a less convincing result and takes roughly 45 to 46 minutes with 15 steps and no lightning preset. The presenter also reports worse results when combining a reference image with the lightning LoRA. These examples give viewers concrete settings and observed limitations for this local AI video workflow.