
Segment Anything 2 (SAM 2) is Meta's open-source model for selecting objects in images and tracking them through video. It's for developers and researchers who need object masks for visual applications or dataset annotation. The model and its web demo can run on your own GPU machine; Meta also provides a hosted demo.
A click, a box, or an existing mask tells SAM 2 which object to select. It can track multiple objects through a video, and additional prompts on any frame let you correct the predicted masks. Its memory of earlier frames helps it follow an object even when it temporarily disappears from view.
The same model handles still images and video. Its streaming architecture processes frames one at a time and is designed for real-time interaction. SAM 2 can also segment unfamiliar objects and visual scenes without training specifically for each one, which matters when your footage differs from the training data.
The Apache 2.0 project includes inference code, downloadable pretrained models, and a web interface you can deploy locally. It uses Python and PyTorch, with model sizes that offer different speed and accuracy tradeoffs. You can also load models through Hugging Face.
For work that needs domain-specific behavior, SAM 2 supports training and fine-tuning on custom image datasets, video datasets, or both. Meta also releases the SA-V video segmentation dataset, with annotations for whole objects, object parts, and objects obscured from view.
Claim this page with an email at ai.meta.com. Segment Anything 2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Segment Anything 2?Promote it
Something wrong or outdated on this page?
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
16.8KUpdated 1 day agoMIT
Docker · Web#Hugging Face integration#Multi-user access#ONNX
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
89Updated 22 hours agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Hugging Face integration#ONNX
PixlStash is an open-source image manager for photographers, AI creators and people curating image datasets. It combines library search with tools for reviewing tags, ranking images and sending selected work through ComfyUI. You can use a desktop app or a self-hosted server with a browser interface.
prodi.gyData Labeling and Annotation
Web#Hugging Face integration#Works offline
5.1KUpdated 1 year agoApache-2.0
Web#Multi-user access#Semantic search
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
CVAT is a browser-based data annotation platform for teams building computer vision datasets. Its open-source Community edition runs on your own infrastructure with Docker and uses the MIT license. CVAT Online is hosted by CVAT, while the Enterprise offering runs in an organization's own cloud or internal environment.
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
Prodigy is a proprietary annotation tool that runs on your own machines, including air-gapped systems without an internet connection. It's for developers and research teams building training and evaluation datasets for custom AI models. The Python library includes a web application where annotators can label data without programming knowledge.