Surya, Qwen3.5 and Mistral OCR: document comparison

Compare OCR results on six document images, including tables and handwriting, plus a Qwen3.5 9B pipeline that scored 69 points in the team's test.

Player not loading? Watch on YouTube

ServerFlow compares specialized OCR models and vision-language models on the same six document images. The tests cover printed text, tables, mathematical formulas, scans and handwriting. The team measures character accuracy against reference text and checks whether each output preserves eight selected fragments of important information.

In this test, dots.mocr leads the individual local models overall and performs particularly well on printed pages and tables. Qwen3.5 4B follows it among local entries and handles difficult material more consistently than the 2B version. The smaller Qwen model nevertheless matches dots.mocr on the table test, with roughly 99% accuracy and all eight selected fragments preserved. Surya OCR 2 performs well on formulas but produces uneven results across documents.

The presenter reports repetition loops or invented text in difficult cases with DeepSeek-OCR 2, Unlimited OCR and FireRed OCR. Mistral OCR's newer cloud version leads the overall comparison, though it sends documents to a third party and charges by volume. These rankings describe this document set, rather than universal performance.

A final offline experiment gives Qwen3.5 9B the original image and drafts from dots.mocr and DeepSeek-OCR 2. The combined pipeline scores 69 points, versus 60 for Qwen alone. Handwriting remains unreliable: high character accuracy can still conceal missing key phrases.