34.4KUpdated 6 months agoMIT
#Batch processing#Multimodal input
CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
34.7KUpdated 2 days agoApache-2.0
Detectron2 is an open-source Python library for developers and researchers building computer vision applications. It provides algorithms for locating objects in images and segmenting image regions, with support for training models and building research projects on top of the library. Facebook AI Research developed it as the successor to Detectron and maskrcnn-benchmark.
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
10.2KUpdated 1 year agoMIT
#Hugging Face integration#Multimodal input
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.
19.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
Segment Anything 2 (SAM 2) is Meta's open-source model for selecting objects in images and tracking them through video. It's for developers and researchers who need object masks for visual applications or dataset annotation. The model and its web demo can run on your own GPU machine; Meta also provides a hosted demo.
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
62.1KUpdated 23 hours agoAGPL-3.0
#ONNX
Ultralytics YOLO is an open-source Python computer vision library for developers building applications that analyze images and video on their own hardware. It supports local and edge deployment, including NVIDIA Jetson, Raspberry Pi and mobile phones. A separate hosted platform provides browser-based annotation, cloud GPU training and managed prediction endpoints.