Grounding DINO + SAM 2: PyQt5 annotation tutorial

Learn to annotate scientific images with text prompts, adjust detection thresholds to 0.05, correct masks by clicking, and export binary PNGs.

Player not loading? Watch on YouTube

This tutorial presents a PyQt5 desktop annotation tool that combines Grounding DINO detection with SAM 2 segmentation. Grounding DINO finds bounding boxes from text descriptions; SAM 2 uses those boxes to create object masks. The application code is a single Python file, and the repository includes download_models.py to place model files where the application expects them. The description also covers downloading models for offline use.

The workflow starts with loading models and an image, then defining classes and detection phrases. In the rat kidney example, the speaker gets no detections at a box threshold of 0.25 and lowers it to 0.05. The mitochondria example shows why threshold changes alone may not suffice: adding "elongated small oval objects" produces detections at 0.15, though some objects still need manual correction. Users can click to add missing masks or delete unwanted ones.

The demonstration exports 22 separate PNG masks with background values of zero and foreground values of one. The speaker explains that these can appear blank in ordinary image viewers. The description also specifies a JSON summary. CPU execution is supported but slower, according to the speaker; the demonstrated workstation has 20 GB of GPU memory. Adjustable interface fonts and saved projects spanning multiple images remain proposed additions.