Nougat OCR tutorial: extract equations and tables from PDFs

Learn Nougat's command-line PDF workflow, MMD output and LaTeX rendering, with examples from a scanned book and an 11-page research paper.

Player not loading? Watch on YouTube

This tutorial, uploaded on September 17, 2023, demonstrates Meta Research's Nougat OCR library on a scanned book excerpt and an 11-page research paper. The source does not specify a Nougat version. Its installation steps and results describe that dated demonstration.

The presenter works in a Colab notebook, explains installation through pip or the source repository, and invokes Nougat from the command line with an input PDF and an output directory. Nougat produces MMD files. When the scanned document triggers a warning about pages skipped due to repetitions, the presenter recommends the optional no-skipping argument.

The second part covers inspecting the output and rendering it as a PDF. The presenter downloads the files and pastes generated content into a LaTeX document in Overleaf, an online compiler. A local LaTeX compiler is also suggested, though the demonstrated OCR workflow runs in Colab.

The presenter judges the extracted equations and tables favorably, but identifies problems with bold headings and spacing around italic text. Figures and graphs are excluded, and the rendered output does not reproduce the original PDF layout exactly. The suggestion that Meta's website examples use extra options or post-processing is the presenter's speculation, rather than an established explanation.