Grounding DINO tutorial: text detection in Google Colab

Learn to load Grounding DINO in Colab and tune detection prompts and thresholds on COCO and microscopy images using Transformers greater than 4.51.0.

Player not loading? Watch on YouTube

This tutorial builds a text-prompted object detection notebook in Google Colab with Grounding DINO base. The speaker loads the model through Hugging Face Transformers, selects a GPU runtime and checks CUDA availability. The stated requirement is Transformers greater than 4.51.0. For repeated use on a local system, he recommends downloading the model weights first and pointing the code to those files.

The notebook defines detection and plotting helpers, then adds non-maximum suppression to reduce overlapping boxes and an area filter to remove oversized detections. The speaker recommends lowercase prompts ending in periods, with multiple descriptions combined into a single string. A COCO image provides an initial test with cats and remote controls before the tutorial moves to kidney H&E and IHC images.

The scientific examples show why prompt wording and thresholds need testing. Changing "glomerulus" to "glomeruli" reduces the detections in the demonstrated IHC image. On a Lucchi electron microscopy slice, the mitochondria prompt returns no detections at 0.25 and two at 0.1. Descriptions such as "a dark oval structure" produce additional boxes, while very low thresholds also pick up background regions. These are results from the examples, rather than evidence of reliable detection across microscopy datasets. The tutorial produces bounding boxes; conversion to masks with SAM is reserved for the next lesson.