Player not loading? Watch on YouTube
This tutorial follows Ostris as he trains an LTX-2.3 character LoRA on footage of himself in one webcam scene. He cuts the recording into 19 clips in Kdenlive, keeps audio enabled, and chooses breaks between words. Most clips are around five seconds. Captions reproduce the spoken wording, including contractions and mistakes; he describes selected gestures to give prompts more control over them.
Training runs in AI Toolkit on RunPod. The speaker suggests a local AI setup using a 5090 with full layer offloading, but demonstrates training on a larger GPU. His configuration uses linear rank 32, cached text embeddings and video latents, automatic frame counts, and low VRAM mode. He starts at 512 and 768 resolution settings with high-noise training, then switches to balanced at 2,000 steps, adds 1024, and finishes at 5,000 steps.
He tests the downloaded LoRA in ComfyUI with the distilled model. At 24 frames per second, he uses 121 frames for five seconds and 241 for ten seconds, then tries 720p and 1080p output. He reports a close likeness and voice match, though names, fine detail, and prompted actions remain imperfect: the hamburger example does not perform the requested bite. The single-scene dataset intentionally ties the character to the room and clothing. His recommendation to train on the base model and generate with the distilled version reflects his experience in this demonstration.
Ostris is the AI Toolkit creator and requests support for his work. His RunPod referral earns credits; the tutorial demonstrates cloud training on a 96 GB GPU, while local RTX 5090 training with full layer offloading is suggested. The 512 and 768 training settings describe approximately square-equivalent pixel budgets, not square video output.