
Depth Anything estimates depth from a single image, making it useful for developers and researchers who need depth maps from ordinary pictures. Its Python implementation runs locally, and a Gradio interface lets you try it on your own machine. An online demo is also available.
The pretrained models estimate relative depth: which parts of a scene are closer or farther away. Separate models fine-tuned for indoor and outdoor scenes estimate metric depth. Small, Base and Large models give users a choice of model size, while the pretrained weights provide a starting point for further fine-tuning.
The model produces finer detail than its predecessor. Compared with Stable Diffusion-based depth estimators such as Marigold and Geowizard, the project reports faster inference, fewer parameters and higher depth accuracy. Its training approach combines labeled synthetic images with real images whose depth labels come from a larger teacher model.
You can process individual images, image collections and video through the repository's scripts. The method estimates depth image by image; larger models provide better consistency between video frames. Hugging Face Transformers offers another way to load the models, with model downloads requiring a Hugging Face connection. Predictions can differ slightly from the repository implementation.
The Small model uses Apache 2.0; Base, Large and Giant models use the non-commercial CC BY-NC 4.0 license. It also includes the DA-2K evaluation benchmark, with sparse depth annotations for testing depth estimation models.
Claim this page and we'll verify you by hand. Depth Anything gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Depth Anything?Promote it
Something wrong or outdated on this page?
135.7KUpdated 1 day agoGPL-3.0
macOS · Windows · Linux · Web#ControlNet#Inpainting#LoRA
ComfyUI is a local visual AI workspace for artists and technical teams who want to control how images, video, audio, 3D models and text are made. Its node canvas shows each model and processing step, so users can build and adjust workflows without writing code. It runs on your hardware.
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
19.9KUpdated 2 years agoApache-2.0
Web#Hugging Face integration
29.9KUpdated 3 weeks ago
macOS · Linux · iOS · Android · Web#Image-to-image#ONNX#PII redaction
90.5KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
1.8KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#Image-to-image#Inpainting
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
Segment Anything 2 (SAM 2) is Meta's open-source model for selecting objects in images and tracking them through video. It's for developers and researchers who need object masks for visual applications or dataset annotation. The model and its web demo can run on your own GPU machine; Meta also provides a hosted demo.
InsightFace is a face analysis toolkit for developers and teams building identity verification, access control, or face editing software. The code uses the MIT license. Its Python tools and self-hosted recognition server run inference on your own hardware. It also offers commercial models and API access for face swapping and deepfake detection.
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
4M is an open-source Python framework for researchers and developers who want one model to handle multiple vision tasks and generate images from mixed inputs. It runs on your own hardware with PyTorch and CUDA. The code uses the Apache 2.0 license, and pretrained model and tokenizer weights are available as safetensors files or through Hugging Face Hub.