
CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.
Its main distinction is zero-shot classification: the same model can compare an image against different sets of text descriptions without retraining. That makes it useful for projects where categories change or a labeled dataset would be costly to build. CLIP learns associations between pictures and language from image-text pairs, rather than learning only a fixed set of object categories.
The Python package exposes image and text features as well as scores for how closely they match. Developers can use those features in their own applications or train a separate classifier on top of the image representations. The available models include ResNet and Vision Transformer architectures. Model loading supports local checkpoints and downloads weights when needed.
CLIP can recognize common objects and has demonstrated capabilities in OCR, video action recognition and geographic classification. Its accuracy depends on the task and the wording of the labels. It struggles with counting objects, estimating distances and distinguishing closely related categories such as car models or flower species. Images unlike those in its training data can also reduce accuracy.
Claim this page with an email at openai.com. CLIP gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find CLIP?Promote it
Something wrong or outdated on this page?
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
10.2KUpdated 1 year agoMIT
#Hugging Face integration#Multimodal input
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.