Favicon of DeepSeek-OCR

DeepSeek-OCR

Local OCR model that converts images and PDFs to text or Markdown. Runs on NVIDIA GPUs with vLLM or Transformers under the MIT license.

DeepSeek-OCR is an open-source OCR model for developers building document processing tools and researchers studying how AI reads text through images. It runs on your own hardware with NVIDIA CUDA GPUs. Its distinctive focus is visual text compression: representing document images with compact sets of vision tokens for a language model to read.

The model can convert document images to Markdown or extract text without layout information. It also accepts prompts to parse figures, describe images, and locate specified text within an image. Those capabilities make it relevant to workflows that need more than a plain transcription, such as extracting document structure or finding a passage on a page.

You can use it through vLLM or Hugging Face Transformers. The supplied vLLM tools support image processing with streaming output, concurrent PDF processing, and batch evaluation. Inference runs locally, so document processing happens on the machine running the model.

It supports fixed image resolutions and a dynamic mode that combines a whole-page view with smaller crops. This gives developers a choice of how much image detail to pass to the model when processing different documents. The project includes Python inference code and uses the MIT license. Its documented GPU environment uses PyTorch and CUDA, with PDF processing demonstrated on an NVIDIA A100.

Similar to DeepSeek-OCR