Favicon of Unstructured

Unstructured

An open-source Python document parser for LLM applications. Run it locally or in Docker to process PDFs, Word files and images under Apache 2.0.

Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.

Its format-specific parsers share an interface with automatic file type detection. OCR extracts text from images and PDFs. Developers can install dependencies for the document formats they need; OCR and PDF processing may require tools such as Tesseract and Poppler. Additional Tesseract language packs support recognition beyond English. The output can feed document preparation and retrieval pipelines.

The local library is separate from Unstructured's hosted services. Batch connectors are provided by the separate unstructured-ingest package, while the hosted Pipelines and Transform products have their own setup and service requirements.

The library sends telemetry by default. Those events include platform details and aggregate processing counts, but exclude document contents, filenames, paths and persistent user or machine identifiers. Users can disable telemetry.

Similar to Unstructured