Favicon of CVAT

CVAT

A data annotation platform for images, video and 3D point clouds, with an MIT-licensed self-hosted edition, automatic labeling and dataset exports.

Screenshot of CVAT website

CVAT is a browser-based data annotation platform for teams building computer vision datasets. Its open-source Community edition runs on your own infrastructure with Docker and uses the MIT license. CVAT Online is hosted by CVAT, while the Enterprise offering runs in an organization's own cloud or internal environment.

Annotators can draw bounding boxes, polygons, masks, keypoints and 3D cuboids, assign tags, and track objects through video. These tools cover detection, segmentation and pose estimation as well as image classification. Community supports images, video and 3D point clouds; the broader platform also supports audio annotation.

Automatic annotation can use your own models or supported models such as Segment Anything, RetinaNet, HRNet and YOLO. Community connects these models through Nuclio and supports frameworks including PyTorch, ONNX, OpenVINO and TensorFlow.

For shared datasets, CVAT divides work into projects, tasks and jobs, with assignments and role-based access. Annotators and reviewers can discuss issues on the annotations. Ground truth, honeypot checks and comparisons between annotators help teams assess label quality, with Community quality checks available through the server API.

CVAT accepts local files and connects to Amazon S3, Azure Blob Storage, Google Cloud Storage and self-hosted storage. It imports and exports formats including COCO, YOLO, Pascal VOC and KITTI. A REST API, Python SDK and command-line tool let developers connect annotation work to their existing data pipelines.

Similar to CVAT