Player not loading? Watch on YouTube
This tutorial fine-tunes Grounding DINO to detect glomeruli in kidney histology images. The presenter uses 24 annotated images: 12 with H&E staining and 12 with IHC staining, with some variation in magnification. He reports that a 20-epoch run took about five minutes on his GPU; training time depends on the dataset and hardware.
The workflow starts in an annotation tool that combines text-prompted Grounding DINO detection with SAM 2 segmentation. The presenter removes incorrect masks and adds missed objects manually, then saves per-image COCO annotations. A merge utility creates training and validation JSON files, using a 20% validation fraction in the demonstration.
The presenter explains that the fine-tuning approach freezes the text and image backbones and updates the cross-modality decoder and detection head. In the training GUI, he selects the base model, supplies the annotation files, chooses a checkpoint folder, and starts training.
The final comparison loads the saved checkpoint back into the annotation tool. On the demonstrated IHC image, the fine-tuned model detects objects at settings where the base model returns none. Some unwanted masks still need manual removal. The example shows an improvement on that image, rather than establishing reliable detection across other tissues or imaging conditions.