Favicon of OpenCLIP

OpenCLIP

A Python library for running and training CLIP image-text models on your own hardware, with local checkpoints and Hugging Face model support.

OpenCLIP is a Python and PyTorch library for developers and researchers who want to match images with text on their own hardware. It implements OpenAI's CLIP approach: images and descriptions become numerical representations that the model can compare. This supports image search and zero-shot classification, where text labels define the categories without a separate classifier trained for each task.

The library provides a shared interface for pretrained models and models you train yourself. Its supported families include original OpenAI CLIP, SigLIP, SigLIP2 and PE, with image encoders such as ViT and ConvNeXt. Pretrained weights include models trained on LAION and DataComp datasets. You can load checkpoints from local files or Hugging Face Hub, so model inference can run locally rather than through a hosted prediction API.

Training support covers fine-tuning and larger runs across multiple GPUs or machines, including SLURM clusters. NVIDIA A100 GPUs are among the hardware used for distributed training. You can combine datasets and adjust their sampling weights, which matters when your training sources differ in size or relevance. The library also supports Hugging Face text encoders and tokenizers.

For applications that need image and text embeddings at scale, clip-retrieval integrates with OpenCLIP. The library includes image preprocessing and text tokenization alongside its model interface, and accepts pretrained checkpoints in local .pt files.

Similar to OpenCLIP