Favicon of Grounding DINO

Grounding DINO

An open-source object detection model that finds objects using text prompts. Run it locally with PyTorch on a CPU or NVIDIA GPU. Apache 2.0 licensed.

Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.

The model takes an image and text together, then returns bounding boxes with matching labels and similarity scores. You can target a particular phrase within a longer description, so a prompt can identify which objects matter rather than ask for every object in the image. Its open-set approach supports detection without training a separate detector for each requested category.

The Python implementation uses PyTorch and includes pretrained GroundingDINO-T and GroundingDINO-B models with Swin-T and Swin-B image backbones. It can run on your own hardware in CPU-only mode or use an NVIDIA GPU through CUDA. The project is open source under the Apache 2.0 license, and it includes Docker support and a Gradio browser interface demo. Hugging Face and Colab also provide hosted demos; those run outside your local machine.

For workflows that need more than bounding boxes, Grounded-SAM combines the detector with Segment Anything for segmentation. Grounded SAM 2 pairs it with SAM 2 for object tracking. Separate example notebooks connect detection to Stable Diffusion and GLIGEN for image editing, using the objects located by the model to guide edits.

Similar to Grounding DINO