
PaddleX is a low-code AI development toolkit for developers building document processing, computer vision and time-series applications on their own hardware. Built on PaddlePaddle, it combines pretrained models with tools for training, inference and deployment. It's open source under Apache 2.0.
Document processing is a substantial part of the toolkit. It supports PaddleOCR-VL and multilingual PP-OCRv5 models, alongside pipelines for reading text, recognizing tables, parsing page layouts and extracting information from documents. Formula recognition and seal text recognition cover more specialized inputs. Document preprocessing can correct orientation and straighten warped page images, while table recognition can return reconstructed tables as HTML.
For image applications, PaddleX includes classification, object detection and segmentation, with specialized capabilities such as small-object detection, image anomaly detection and human keypoint detection. Its time-series tools cover forecasting and anomaly detection. Developers can use complete pipelines or combine individual model modules, and the development tools include both Python interfaces and a graphical interface.
Models can run locally on CPUs or NVIDIA GPUs, with support for Kunlunxin, Ascend and Cambricon hardware as well. Linux and Windows are supported. Deployment choices include self-hosted services through Docker and execution on edge devices. Paddle Inference and ONNX Runtime are available as inference backends, and pipeline benchmarks measure total processing time as well as time spent in individual modules.
Claim this page and we'll verify you by hand. PaddleX gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PaddleX?Promote it
Something wrong or outdated on this page?
6.4KUpdated 22 hours agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
9.4KUpdated 1 day agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
15.5KUpdated 3 hours agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
1.6KUpdated 2 weeks agoApache-2.0
macOS · Windows · Linux · Docker#MCP#Multimodal input#Ollama integration
ExtractThinker is a Python library for developers who need structured data from documents, such as invoice fields their application can use. It pairs document parsers with an LLM and returns results that follow a Pydantic schema. It's open source under Apache 2.0 and runs on macOS, Windows and Linux.
29.5KUpdated 21 hours agoApache-2.0
macOS · Linux · Docker#Guardrails#Multilingual#ONNX
OpenDataLoader PDF is a local PDF parser for developers building document search and RAG systems, with free accessibility auditing and auto-tagging. It converts PDFs into Markdown, JSON and HTML while preserving the structure needed to retrieve passages and cite their source. The open-source parser uses the Apache 2.0 license and runs locally without a GPU or API keys.