Player not loading? Watch on YouTube
This tutorial builds a text-prompted object detection notebook in Google Colab with Grounding DINO base. The speaker loads the model through Hugging Face Transformers, selects a GPU runtime and checks CUDA availability. The stated requirement is Transformers greater than 4.51.0. For repeated use on a local system, he recommends downloading the model weights first and pointing the code to those files.
The notebook defines detection and plotting helpers, then adds non-maximum suppression to reduce overlapping boxes and an area filter to remove oversized detections. The speaker recommends lowercase prompts ending in periods, with multiple descriptions combined into a single string. A COCO image provides an initial test with cats and remote controls before the tutorial moves to kidney H&E and IHC images.
The scientific examples show why prompt wording and thresholds need testing. Changing "glomerulus" to "glomeruli" reduces the detections in the demonstrated IHC image. On a Lucchi electron microscopy slice, the mitochondria prompt returns no detections at 0.25 and two at 0.1. Descriptions such as "a dark oval structure" produce additional boxes, while very low thresholds also pick up background regions. These are results from the examples, rather than evidence of reliable detection across microscopy datasets. The tutorial produces bounding boxes; conversion to masks with SAM is reserved for the next lesson.