Open Models for Detection and Segmentation

Open models for spotting objects, segmenting images and captioning them, including YOLO, Segment Anything 2 and CLIP.

10 tools
An open-source image and text model for local image classification without task-specific training. Runs through PyTorch on CPU or CUDA GPUs under MIT.

34.4KUpdated 6 months agoMIT

#Batch processing#Multimodal input

An open-source object detection model that finds objects using text prompts. Run it locally with PyTorch on a CPU or NVIDIA GPU. Apache 2.0 licensed.

10.6KUpdated 2 years agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input

Open-source Python library for object detection and segmentation, with pretrained models and deployment exports. Licensed under Apache 2.0.

34.7KUpdated 2 days agoApache-2.0

Face analysis toolkit for self-hosted recognition on CPU or NVIDIA GPU, local video face redaction, and commercially licensed models.

29.9KUpdated 3 weeks ago

macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#Quantization

An open-source vision model that runs locally with PyTorch and Hugging Face Transformers, supports CPU or CUDA GPUs, and uses the MIT license.

huggingface.coComputer Vision Models

#Hugging Face integration#Multimodal input#Structured output

An open-weight vision model for image questions, captions and object detection. Run it locally or use hosted inference and fine-tuning.

10.1KUpdated 5 months agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline

Open-source vision-language models for visual chat, document questions and image retrieval, with downloadable weights and Hugging Face Transformers support.

10.2KUpdated 1 year agoMIT

#Hugging Face integration#Multimodal input

An open-source segmentation model for images and videos. Run it on your own GPU machine, track objects across frames, and refine masks with prompts.

19.9KUpdated 2 years agoApache-2.0

Web#Hugging Face integration

Favicon of PaddleOCR

PaddleOCR

2 videos
An open source OCR toolkit that runs locally, converts PDFs and images to Markdown or JSON, and supports multilingual text under Apache 2.0.

90.4KUpdated 2 weeks agoApache-2.0

Web#Multilingual#ONNX#Structured output

Open-source computer vision library for local detection, segmentation and tracking, with AGPL-3.0 licensing and exports to ONNX, TensorRT and CoreML.

62.1KUpdated 23 hours agoAGPL-3.0

#ONNX

More in Open Models