Nemotron post-training: data blends and tool calls

Learn how NVIDIA approaches Nemotron post-training, including data blends, tool-call penalties and published recipes that require your own compute.

Player not loading? Watch on YouTube

NVIDIA researchers explain how they approach Nemotron post-training, with an emphasis on instruction following and AI agent capabilities. They describe supervised fine-tuning (SFT) as training on structured conversations, tool responses and task prompts rather than the unstructured text used in pre-training.

The discussion focuses on how training data affects behavior. The team tests subsets in isolation before combining them and examines trajectory lengths, tool diversity and correlations. One example is a model learning to call a tool whenever it appears in the prompt, even when the task does not need it. The researchers describe mixing general chat data into specialized training and using multiple expert teachers through on-policy distillation to help preserve broader capabilities.

For tool-call problems, Christian describes penalties for malformed calls, synthetic recovery examples created with NeMo Data Designer and length penalties intended to encourage efficiency. Checkpoint selection also includes benchmarks, safety review and practical testing in chat interfaces and coding harnesses.

The session points to published datasets and training recipes. Reproducing the full workflow requires sufficient compute, and the researchers caution that a data blend can behave differently depending on the base model. This is a methods discussion rather than a step-by-step installation guide.