TensorRT Model Connect: voice and video deployment tutorial

Learn TensorRT Model Connect's build-and-run workflow, with Nemotron Voice chat on an RTX 5090 and Cosmos 3 video generation across two DGX Sparks.

Player not loading? Watch on YouTube

TensorRT Model Connect is presented as a feature on top of TensorRT for deploying supported model families through a C++ runtime. The walkthrough starts with a small text model: a Python command builds a deployable bundle, then a command-line utility runs inference with a prompt and chat template. The speaker describes graph fusion and tactic selection during compilation, with build time varying by model size.

The examples extend beyond a local LLM to speech transcription, transcript refinement, reasoning, semantic search, image segmentation and depth estimation. The presenters say the single-device demos use an RTX 5090. Nemotron Voice demonstrates interruptible audio conversation through a single model, rather than a separate speech recognition, language model and speech synthesis pipeline.

For Cosmos 3 video generation, the tutorial uses two DGX Sparks connected by Ethernet. It outlines package installation, passwordless SSH, a Docker image build and a prepare command before inference with context parallelism across both devices.

The stated limits matter for anyone planning to run models locally: Model Connect handles inference rather than training, the model architecture must be supported, and its weights must fit the available GPU memory. Unsupported quantization requests can fail the build. The presenters also say some demonstrated applications would reach the repository after the stream.