NVIDIA Cosmos 3 overview and Nano inference demo

Learn how Cosmos 3 combines reasoning and generation, with a Nano text-to-video demo on DGX Spark and examples of object consistency problems.

Player not loading? Watch on YouTube

NVIDIA Cosmos 3 is the subject of this overview of world foundation models for physical AI. The speaker describes a model that accepts and generates five modalities: text, images, video, audio and actions. Robotics examples focus on generating synthetic training data and predicting what happens after an action.

The architecture discussion explains a two-tower mixture of transformers. An autoregressive reasoner handles inputs, while a diffusion tower generates outputs. The speaker gives Cosmos 3 Super's size as 32 billion parameters per tower and Nano's as 8 billion per tower, interpreting Nano as roughly 16 billion parameters overall. An edge version intended for on-device use had not been released at the time of the recording.

The speaker reports testing Nano inference on a DGX Spark. The demo uses a JSON file to select text-to-video mode and describe robotic arms picking fruit from a wooden display cabinet. He notes occasional consistency problems when objects drop. Other discussed modes include image-to-video, video-to-video and forward dynamics, which predicts the next state given an action.

The overview also covers released open weights and a technical report documenting architecture, pretraining and supervised fine-tuning data. It provides context for local AI experimentation with robotics; the demo does not establish minimum hardware requirements or generation speed.