OpenCLIP is a Python and PyTorch library for developers and researchers who want to match images with text on their own hardware. It implements OpenAI's CLIP approach: images and descriptions become numerical representations that the model can compare. This supports image search and zero-shot classification, where text labels define the categories without a separate classifier trained for each task.
The library provides a shared interface for pretrained models and models you train yourself. Its supported families include original OpenAI CLIP, SigLIP, SigLIP2 and PE, with image encoders such as ViT and ConvNeXt. Pretrained weights include models trained on LAION and DataComp datasets. You can load checkpoints from local files or Hugging Face Hub, so model inference can run locally rather than through a hosted prediction API.
Training support covers fine-tuning and larger runs across multiple GPUs or machines, including SLURM clusters. NVIDIA A100 GPUs are among the hardware used for distributed training. You can combine datasets and adjust their sampling weights, which matters when your training sources differ in size or relevance. The library also supports Hugging Face text encoders and tokenizers.
For applications that need image and text embeddings at scale, clip-retrieval integrates with OpenCLIP. The library includes image preprocessing and text tokenization alongside its model interface, and accepts pretrained checkpoints in local .pt files.
Claim this page and we'll verify you by hand. OpenCLIP gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OpenCLIP?Promote it
Something wrong or outdated on this page?
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
2.8KUpdated 1 month agoMIT
macOS#Batch processing#Hugging Face integration#LoRA
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
937Updated 2 years agoApache-2.0
#Hugging Face integration#Multimodal input#Works offline
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
ColPali is a local AI document retrieval library for developers and researchers building document search or retrieval-augmented generation systems. It searches pages as images, using their text, charts and layout together rather than relying on a separate OCR pipeline. The colpali-engine package is deprecated; its maintainers recommend Sentence Transformers for new projects and production use.
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
Molmo is Ai2's family of vision-language models, with code for running and training models on your own hardware. It's for developers and researchers who need to work with images and text, adapt a model, or evaluate it against visual tasks. The Python codebase is open source under Apache 2.0 and builds on OLMo, adding image encoding and generative evaluation.
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.