
NVIDIA Cosmos is a platform of world models for developers building robots, autonomous vehicles, and industrial vision systems. Its downloadable models run on your own NVIDIA hardware; NVIDIA also offers a hosted catalog for browser access. Cosmos combines visual reasoning, simulation, and action generation in a shared Mixture-of-Transformers architecture.
The models can interpret objects, interactions, and intent in images or video, and generate possible future scenes from text, visual input, sound, and actions. Developers can use those outputs as synthetic training data or compare behaviors in physics-grounded simulations before testing them in the real world. Robot policy learning supports adaptation to specific tasks, camera layouts, and robot embodiments. Video analysis covers live feeds and recorded footage, including contextual alerts, captioning, and search through NVIDIA Metropolis Blueprint.
Hardware needs vary by model. Cosmos3-Edge supports Jetson AGX Orin and Thor, while Nano and Super target workstation or data-center GPUs such as RTX PRO 6000, H100, H200, and B200. Supported backends include Diffusers, Transformers, vLLM, TensorRT-LLM, and NIM, with OpenAI-compatible API serving available.
Cosmos Framework provides training and optimization tools. Cosmos Curator filters, annotates, and deduplicates sensor data, while Cosmos Evaluator scores generated outputs. The video quickstart requires approved Hugging Face access to a guardrail model. Generation uses guardrails by default, but outputs can still contain inconsistent motion, changing object shapes, or implausible physics; safety-critical applications need additional validation.
Claim this page with an email at nvidia.com. NVIDIA Cosmos gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find NVIDIA Cosmos?Promote it
Something wrong or outdated on this page?
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
25.1KUpdated 2 years agoApache-2.0
macOS · Web#LoRA#Multimodal input#Quantization
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
LLaVA is a family of vision-language models for researchers and developers who want to ask questions about images on their own hardware. It pairs a CLIP vision encoder with a language model to support image descriptions, visual reasoning and reading text in pictures. Its Python code is open source under Apache 2.0; the project places research-use restrictions on its data and checkpoints, with additional terms from the underlying models.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.