Favicon of Segment Anything 2

Segment Anything 2

An open-source segmentation model for images and videos. Run it on your own GPU machine, track objects across frames, and refine masks with prompts.

Screenshot of Segment Anything 2 website

Segment Anything 2 (SAM 2) is Meta's open-source model for selecting objects in images and tracking them through video. It's for developers and researchers who need object masks for visual applications or dataset annotation. The model and its web demo can run on your own GPU machine; Meta also provides a hosted demo.

A click, a box, or an existing mask tells SAM 2 which object to select. It can track multiple objects through a video, and additional prompts on any frame let you correct the predicted masks. Its memory of earlier frames helps it follow an object even when it temporarily disappears from view.

The same model handles still images and video. Its streaming architecture processes frames one at a time and is designed for real-time interaction. SAM 2 can also segment unfamiliar objects and visual scenes without training specifically for each one, which matters when your footage differs from the training data.

The Apache 2.0 project includes inference code, downloadable pretrained models, and a web interface you can deploy locally. It uses Python and PyTorch, with model sizes that offer different speed and accuracy tradeoffs. You can also load models through Hugging Face.

For work that needs domain-specific behavior, SAM 2 supports training and fine-tuning on custom image datasets, video datasets, or both. Meta also releases the SA-V video segmentation dataset, with annotations for whole objects, object parts, and objects obscured from view.

Similar to Segment Anything 2