
CVAT is a browser-based data annotation platform for teams building computer vision datasets. Its open-source Community edition runs on your own infrastructure with Docker and uses the MIT license. CVAT Online is hosted by CVAT, while the Enterprise offering runs in an organization's own cloud or internal environment.
Annotators can draw bounding boxes, polygons, masks, keypoints and 3D cuboids, assign tags, and track objects through video. These tools cover detection, segmentation and pose estimation as well as image classification. Community supports images, video and 3D point clouds; the broader platform also supports audio annotation.
Automatic annotation can use your own models or supported models such as Segment Anything, RetinaNet, HRNet and YOLO. Community connects these models through Nuclio and supports frameworks including PyTorch, ONNX, OpenVINO and TensorFlow.
For shared datasets, CVAT divides work into projects, tasks and jobs, with assignments and role-based access. Annotators and reviewers can discuss issues on the annotations. Ground truth, honeypot checks and comparisons between annotators help teams assess label quality, with Community quality checks available through the server API.
CVAT accepts local files and connects to Amazon S3, Azure Blob Storage, Google Cloud Storage and self-hosted storage. It imports and exports formats including COCO, YOLO, Pascal VOC and KITTI. A REST API, Python SDK and command-line tool let developers connect annotation work to their existing data pipelines.
Claim this page with an email at cvat.ai. CVAT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find CVAT?Promote it
Something wrong or outdated on this page?
89Updated 22 hours agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Hugging Face integration#ONNX
PixlStash is an open-source image manager for photographers, AI creators and people curating image datasets. It combines library search with tools for reviewing tags, ranking images and sending selected work through ComfyUI. You can use a desktop app or a self-hosted server with a browser interface.
10.8KUpdated 8 months agoMIT
Windows · Docker · Web#Multi-user access#Multilingual
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
3.3KUpdated 2 weeks agoApache-2.0
Docker · Web#LLM tracing#MCP
Laminar is an open-source platform for developers who need to see why an AI agent failed and check whether a fix worked. You can self-host it with Docker or on Kubernetes, including AWS and GCP, or use its managed cloud service. It uses the Apache 2.0 license.
5.1KUpdated 1 year agoApache-2.0
Web#Multi-user access#Semantic search
Argilla is an open-source data annotation and feedback tool for AI engineers and domain experts who build training and evaluation datasets. You can run your own Argilla server or deploy it on Hugging Face Spaces. It's licensed under Apache 2.0.
1.3KUpdated 7 months agoApache-2.0
Windows · Docker#Hugging Face integration#Multimodal input#OpenAI-compatible API
Doccano is a self-hosted text annotation tool for machine learning practitioners who need labeled training or evaluation data. It runs on your own machine or server, with a browser interface and Docker support. The software is open source under the MIT license.
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.