Favicon of Nemotron

Nemotron

A family of AI models you can run offline with Ollama, llama.cpp or LM Studio, with open weights and training data for building specialized agents.

Screenshot of Nemotron website

Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.

The family covers different hardware and workloads. Nano targets PCs and edge devices; Lightning handles frequent, specialized tasks in long-running agents. Super supports deployment on a single data center GPU, while Ultra targets multi-GPU systems for complex planning and reasoning. Nano Omni understands text, images, video and audio within one model.

Local options include Ollama, LM Studio and llama.cpp, with GGUF models available through Hugging Face. Server deployments support vLLM, SGLang and TensorRT-LLM on NVIDIA GPUs. NVIDIA NIM provides deployable inference microservices, while OpenRouter and managed inference providers offer hosted access. Those hosted services run inference outside your hardware.

Beyond general reasoning, Nemotron includes retrieval models for embeddings and reranking, plus document parsing that extracts text and tables from complex layouts. Speech models cover transcription, speech generation and translation. Safety models handle multilingual moderation, jailbreak detection and personal information detection.

The Apache 2.0 developer repository includes training and fine-tuning pipelines, reinforcement learning recipes and examples for retrieval and tool-using agents. Model weights have release-specific terms: for example, Llama-3.3-Nemotron-Super-49B-v1 uses the NVIDIA Open Model License with additional Llama terms. Check the chosen model card before deployment.

Similar to Nemotron