Favicon of ExtractThinker

ExtractThinker

Open-source Python document extraction library using Pydantic schemas and LLMs, with Ollama support and an Apache 2.0 license.

Screenshot of ExtractThinker website

ExtractThinker is a Python library for developers who need structured data from documents, such as invoice fields their application can use. It pairs document parsers with an LLM and returns results that follow a Pydantic schema. It's open source under Apache 2.0 and runs on macOS, Windows and Linux.

The schema defines both the fields to extract and the rules those values must satisfy. You can add application-specific validation, and extraction raises an error when validation fails rather than inventing defaults for required fields. A valid result can still contain factual mistakes, so schema checks don't replace evaluation on your own documents.

Document loaders cover PDFs, images, tables and spreadsheets. Classification and splitting help process bundles that contain different document types. Text PDFs can use pypdf; scans need OCR or a vision-capable model, while image extraction requires a loader that supplies page images.

The library supports Ollama for local LLM use as well as providers such as OpenAI, Anthropic and Groq. Document loading can run without provider credentials or a model call. Extraction sends document content to the configured model, so a cloud provider receives that content when you choose a hosted backend. Provider access requires the relevant credentials and may carry charges.

An optional MCP service exposes the library through Docker or a native Python process.

Similar to ExtractThinker