Player not loading? Watch on YouTube
This Russian-language presentation is available with English captions. Serverflow compares five OCR models with Qwen3.6-35B, a vision-language model, for local AI document processing. The lineup includes MinerU2.5-Pro-1.2B, PaddleOCR-VL-1.5, GLM-OCR, Chandra-OCR-2 and olmOCR-2-7B. The tests examine extracted text, Markdown structure and whether models preserve important information for later search or RAG processing.
The team measures text accuracy using character error rate and checks eight selected fragments per document with a unit-test pass rate. These checks matter because a readable result can still omit a name, amount or key phrase. Bounding boxes also show how some models divide pages into document elements; olmOCR has no bounding-box visualization in this run.
The speaker reports little difference on the simple document page. MinerU performs well on tables and structured PDFs despite its 1.2 billion parameters. PaddleOCR is presented as a compact option for basic recognition, with weaker results on complex material. Qwen and olmOCR perform better in difficult cases, while Qwen requires more video memory and heavier inference, according to the speaker.
Handwriting remains unreliable across the tested models. The team recommends manual verification and checking whether essential information survives, rather than relying on character accuracy alone. These conclusions describe this test set, not a universal ranking.