Favicon of Florence-2

Florence-2

An open-source vision model that runs locally with PyTorch and Hugging Face Transformers, supports CPU or CUDA GPUs, and uses the MIT license.

Screenshot of Florence-2 website

Florence-2 is Microsoft's open-source vision model for developers who want to process images on their own hardware. It handles several image tasks through text prompts, so one model can generate descriptions, read text and locate objects. It runs locally with PyTorch and Hugging Face Transformers on a CPU or CUDA GPU, and uses the MIT license.

Captioning ranges from short descriptions to detailed accounts of an image. For applications that need locations as well as words, object detection returns labels and bounding boxes, with optional confidence scores. Dense region captioning describes individual parts of an image, while phrase grounding connects words in a supplied caption to the areas they describe. The model also supports segmentation and region proposals.

OCR extracts text from images. It can return the text alone or include coordinates for each text region, which is useful when an application needs to preserve the connection between words and their position in a picture.

Its main distinction is the shared model across these tasks, rather than a separate specialist model for each one. Florence-2 uses an architecture that generates output sequences from image and text inputs, and supports both use without task-specific training and fine-tuning. The model family includes base and large variants, plus fine-tuned versions trained across multiple downstream tasks. Hugging Face distributes the model weights in Safetensors format.

Similar to Florence-2