Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
The model takes an image and text together, then returns bounding boxes with matching labels and similarity scores. You can target a particular phrase within a longer description, so a prompt can identify which objects matter rather than ask for every object in the image. Its open-set approach supports detection without training a separate detector for each requested category.
The Python implementation uses PyTorch and includes pretrained GroundingDINO-T and GroundingDINO-B models with Swin-T and Swin-B image backbones. It can run on your own hardware in CPU-only mode or use an NVIDIA GPU through CUDA. The project is open source under the Apache 2.0 license, and it includes Docker support and a Gradio browser interface demo. Hugging Face and Colab also provide hosted demos; those run outside your local machine.
For workflows that need more than bounding boxes, Grounded-SAM combines the detector with Segment Anything for segmentation. Grounded SAM 2 pairs it with SAM 2 for object tracking. Separate example notebooks connect detection to Stable Diffusion and GLIGEN for image editing, using the objects located by the model to guide edits.
Claim this page and we'll verify you by hand. Grounding DINO gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Grounding DINO?Promote it
Something wrong or outdated on this page?
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
huggingface.coComputer Vision Models
#Hugging Face integration#Multimodal input#Structured output
10.2KUpdated 1 year agoMIT
#Hugging Face integration#Multimodal input
10.1KUpdated 5 months agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#Works offline
34.4KUpdated 6 months agoMIT
#Batch processing#Multimodal input
CLIP is an open-source image and text model that lets developers and researchers classify images using labels written in natural language, without collecting training examples for each task. It runs locally through PyTorch on a CPU or CUDA GPU. The code and model weights use the MIT license.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.
Moondream is a vision model for developers building software that needs to understand images. It can answer questions about a picture, write captions, locate objects, identify points and segment regions. The open-weight models can run on your own hardware, including in an air-gapped environment. The repository code is licensed under Apache 2.0; check each model checkpoint’s own terms for use.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.