
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
The family covers different hardware and workloads. Nano targets PCs and edge devices; Lightning handles frequent, specialized tasks in long-running agents. Super supports deployment on a single data center GPU, while Ultra targets multi-GPU systems for complex planning and reasoning. Nano Omni understands text, images, video and audio within one model.
Local options include Ollama, LM Studio and llama.cpp, with GGUF models available through Hugging Face. Server deployments support vLLM, SGLang and TensorRT-LLM on NVIDIA GPUs. NVIDIA NIM provides deployable inference microservices, while OpenRouter and managed inference providers offer hosted access. Those hosted services run inference outside your hardware.
Beyond general reasoning, Nemotron includes retrieval models for embeddings and reranking, plus document parsing that extracts text and tables from complex layouts. Speech models cover transcription, speech generation and translation. Safety models handle multilingual moderation, jailbreak detection and personal information detection.
The Apache 2.0 developer repository includes training and fine-tuning pipelines, reinforcement learning recipes and examples for retrieval and tool-using agents. Model weights have release-specific terms: for example, Llama-3.3-Nemotron-Super-49B-v1 uses the NVIDIA Open Model License with additional Llama terms. Check the chosen model card before deployment.
Claim this page with an email at developer.nvidia.com. Nemotron gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Nemotron?Promote it
Something wrong or outdated on this page?
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
25KUpdated 2 years agoApache-2.0
macOS · Web#LoRA#Multimodal input#Quantization
1.9KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#llama.cpp backend#Quantization
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
7.7KUpdated 12 months ago
#Hugging Face integration#Multimodal input#Quantization
Llama is Meta's family of large language models for developers, researchers and businesses that want to run models on their own hardware or servers. Its downloadable weights let you build generative AI applications with local inference. Access requires license acceptance and approval, and the weights use custom licensing for research and commercial use.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
LLaVA is a family of vision-language models for researchers and developers who want to ask questions about images on their own hardware. It pairs a CLIP vision encoder with a language model to support image descriptions, visual reasoning and reading text in pictures. Its Python code is open source under Apache 2.0; the project places research-use restrictions on its data and checkpoints, with additional terms from the underlying models.
Augmentoolkit turns your documents into training data for a custom LLM that learns a particular subject. It's for researchers, developers and hobbyists who want models trained on their own material, such as research papers or fictional lore. The Python toolkit is open source under the MIT license and runs on macOS and Linux, with WSL recommended for Windows.