Player not loading? Watch on YouTube
NVIDIA's presenters introduce Cosmos 3 as a foundation model for physical AI, with text, image, video, audio and action inputs. They explain how an autoregressive reasoner and a diffusion generator share attention. Nano has 16 billion parameters, split equally between the towers; Super has about 64 billion. Developers can also call the reasoning tower independently for vision-language tasks.
The repository walkthrough explains where to find inference scripts and post-training guides to run models locally. A JSON input selects the generation mode and supplies the prompt. Model downloads require a Hugging Face token and acceptance of the license. The speakers describe the model and supporting materials as available under MDW 1.1. They show documentation and example commands but do not execute inference or training during the session.
Cosmos Framework handles inference and post-training. Curator prepares and captions video training data, while Evaluator checks generated data for problems such as hallucinations and inconsistent objects. The action section distinguishes forward dynamics, which predicts video from actions, inverse dynamics, which infers actions from video, and policy generation.
The presenters report testing training on H100 GPUs and say DGX Spark is for inference rather than the demonstrated training steps. They caution that action latency depends on hardware and model size and can take seconds. Sensor outputs are not currently supported, and drift requires ongoing evaluation.