Player not loading? Watch on YouTube
This Cosmos Labs tutorial explains how NVIDIA TAO agent skills let a coding assistant prepare and orchestrate Cosmos 3 post-training through natural language prompts. The example uses Cosmos 3 Nano and the Woven traffic safety dataset, which contains one-minute videos with questions and answers.
The presenters compare full supervised fine-tuning with parameter-efficient methods such as LoRA, then explain how AutoML searches training hyperparameters. In the Codex walkthrough, the agent checks hardware, drivers and Docker compatibility, prepares baseline evaluation and training configurations, and waits for approval. The live demonstration stops before training; the presenter uses previously completed experiment results to explain the outcome.
NVIDIA reports exact-match accuracy rising from 54.41% to about 87% after one LoRA run, then to 93.35% with AutoML. The reported sweep ran 43 trials over about 19.5 hours on four GPU nodes. These figures describe this experiment, rather than an expected gain for every dataset or machine.
Setup requires credentials for Hugging Face model downloads, NGC containers and an LLM endpoint used by AutoML to generate configurations. The presenters discuss local execution, Brev and Slurm targets. They say the skills can run through a coding agent on Mac or Windows, while the training or deployment target needs a GPU. Deployment is described as another agent prompt, without a deployment walkthrough.