Player not loading? Watch on YouTube
This sponsored tutorial follows a custom YOLO26 segmentation project through image collection, annotation, training and PC deployment. It uses Ultralytics Platform for the uploaded dataset and training, then downloads the model for local inference with a Python webcam script.
The speaker recommends about 50 to 100 images for a proof of concept, captured with the camera angle and lighting expected in the application. Annotation starts with manual outlines before moving to SAM 3, which proposes object masks from clicks. Users still need to check class assignments and correct outlines. The platform also accepts a trained model to propose labels.
Training covers model size, epochs, resolution and GPU selection. The speaker suggests YOLO26S-seg as a starting point and explains the speed and accuracy tradeoffs. The walkthrough uses Anaconda to create a Python 3.13 environment and installs the Ultralytics package. An optional CUDA-enabled PyTorch installation targets PCs with Nvidia graphics cards. The downloaded model runs on a USB webcam; the speaker reports roughly 25 to 30 FPS in this demonstration and notes difficulty detecting a carrot.
The final example, VeggieVision, uses segmentation masks to estimate produce dimensions, volume and weight through spherical estimation or disk integration. Its GitHub code provides an example of how to work with mask coordinates.