Unlimited-OCR vs DeepSeek-OCR: long document review

Learn how Unlimited-OCR handles long PDFs, why its KV cache stays flat, and how the 32,000-token input limit affects single-pass parsing.

Player not loading? Watch on YouTube

The speaker tests Baidu's Unlimited-OCR with a roughly 50-page, 5 MB PDF in a custom app built with Claude Code. The demo streams extracted text and page coordinates. The speaker describes the open-weights model as a DeepSeek-OCR fine-tune with just over 3 billion parameters and about half a billion active.

The review separates traditional OCR, structure-aware parsers and vision language models. It explains DeepSeek-OCR's optical compression, which reduces pages to fewer vision tokens at the cost of some accuracy. Unlimited-OCR supports two resolution modes: Gundam processes high-resolution pages individually, while base mode accepts a whole document at lower resolution.

Reference Sliding Window Attention keeps the image tokens available while retaining only a short window of generated text. The speaker attributes flat cache memory and steady generation speed to this design, citing research that reports speed gains of up to 35% for inputs above 6,000 image tokens.

The limits matter for local AI document processing. Input tops out at 32,000 tokens, and the speaker questions the benchmark comparison because it omits newer models. Their usual choice is Docling, with local VLMs for harder documents. They favor parallel page processing for most ingestion work, but see a use for sequential parsing when tables or sentences cross page breaks.