NVIDIA Cosmos 3: post-training for robot action prediction

Learn to post-train Cosmos 3 with YAM robot data, configure joint action modes, convert checkpoints, and compare predicted video with ground truth.

Player not loading? Watch on YouTube

This tutorial explains how to post-train NVIDIA Cosmos 3 with robot recordings for action-conditioned prediction. The speaker describes three modes: forward dynamics generates future video from an action trajectory, a conditioning frame and a task description; inverse dynamics recovers actions from video and a task; policy mode predicts actions and video from a frame and a task.

The example uses a YAM robot pick-and-place dataset in LeRobot V3 format. A robot-specific adapter turns seven-channel recordings into video-action samples and converts absolute targets into relative targets. The Python experiment config connects the dataset to the training components. It sets the embodiment to YAM, selects joint mode to train all three modes together, and points to the dataset adapter.

The workflow converts the starting model to DCP format for distributed training, then converts the trained checkpoint to safetensors for inference. The speaker says training takes several hours. An evaluation script runs all three modes, and the final forward-dynamics example places generated video beside ground truth.

The supplied guide describes four GPUs on a GB200 node for this tutorial and a GitHub guide tested on eight H100 GPUs. It directs users to adjust process count and data parallelism for their hardware. This is a multi-GPU robotics training workflow; the example comparison does not establish prediction accuracy across other tasks or robots.