Favicon of CLIP

CLIP

An open-source image and text model for local image classification without task-specific training. Runs through PyTorch on CPU or CUDA GPUs under MIT.

Screenshot of CLIP website

CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.

Its main distinction is zero-shot classification: the same model can compare an image against different sets of text descriptions without retraining. That makes it useful for projects where categories change or a labeled dataset would be costly to build. CLIP learns associations between pictures and language from image-text pairs, rather than learning only a fixed set of object categories.

The Python package exposes image and text features as well as scores for how closely they match. Developers can use those features in their own applications or train a separate classifier on top of the image representations. The available models include ResNet and Vision Transformer architectures. Model loading supports local checkpoints and downloads weights when needed.

CLIP can recognize common objects and has demonstrated capabilities in OCR, video action recognition and geographic classification. Its accuracy depends on the task and the wording of the labels. It struggles with counting objects, estimating distances and distinguishing closely related categories such as car models or flower species. Images unlike those in its training data can also reduce accuracy.

Similar to CLIP