
Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.
Captioning ranges from short descriptions to detailed accounts of an image. For applications that need locations as well as words, object detection returns labels and bounding boxes, with optional confidence scores. Dense region captioning describes individual parts of an image, while phrase grounding connects words in a supplied caption to the areas they describe. The model also supports segmentation and region proposals.
OCR extracts text from images. It can return the text alone or include coordinates for each text region, which is useful when an application needs to preserve the connection between words and their position in a picture.
Its main distinction is the shared model across these tasks, rather than a separate specialist model for each one. Florence-2 uses an architecture that generates output sequences from image and text inputs, and supports both use without task-specific training and fine-tuning. The model family includes base and large variants, plus fine-tuned versions trained across multiple downstream tasks. Hugging Face distributes the model weights in Safetensors format.
Claim this page and we'll verify you by hand. Florence-2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Florence-2?Promote it
Something wrong or outdated on this page?
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
huggingface.coOCR and Document Scanning
Linux#Batch processing#Hugging Face integration#Multimodal input
23.9KUpdated 8 months agoMIT
Linux#Batch processing#Hugging Face integration#Multimodal input
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
10.2KUpdated 1 year agoMIT
#Hugging Face integration#Multimodal input
Qwen2.5-VL is a vision-language model you can run on your own hardware to answer questions about images and video. It's aimed at developers building document processing tools, visual assistants and agents that interact with computer or phone screens. The instruction-tuned 7B model has Apache 2.0 licensing and works with Hugging Face Transformers, with weights available in Safetensors format.
DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
InternVL is a family of downloadable vision-language models for developers and researchers building AI that can interpret images and discuss them in text. It combines visual recognition with language models, supporting both multimodal chat and tasks such as image classification and image-text retrieval.