Depth Anything

A local AI depth estimation model that processes images and video, with a Python implementation, Hugging Face Transformers support and a local Gradio demo.

Screenshot of Depth Anything website

Depth Anything estimates depth from a single image, making it useful for developers and researchers who need depth maps from ordinary pictures. Its Python implementation runs locally, and a Gradio interface lets you try it on your own machine. An online demo is also available.

The pretrained models estimate relative depth: which parts of a scene are closer or farther away. Separate models fine-tuned for indoor and outdoor scenes estimate metric depth. Small, Base and Large models give users a choice of model size, while the pretrained weights provide a starting point for further fine-tuning.

The model produces finer detail than its predecessor. Compared with Stable Diffusion-based depth estimators such as Marigold and Geowizard, the project reports faster inference, fewer parameters and higher depth accuracy. Its training approach combines labeled synthetic images with real images whose depth labels come from a larger teacher model.

You can process individual images, image collections and video through the repository's scripts. The method estimates depth image by image; larger models provide better consistency between video frames. Hugging Face Transformers offers another way to load the models, with model downloads requiring a Hugging Face connection. Predictions can differ slightly from the repository implementation.

The Small model uses Apache 2.0; Base, Large and Giant models use the non-commercial CC BY-NC 4.0 license. It also includes the DA-2K evaluation benchmark, with sparse depth annotations for testing depth estimation models.

Similar to Depth Anything