Favicon of xberg (formerly Kreuzberg)

xberg (formerly Kreuzberg)

An open-source document extraction engine with a Rust core, CPU-only processing, Docker deployment, and support for Ollama, LM Studio, and vLLM.

xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.

The Rust core runs on CPU without requiring a GPU. You can embed it through language bindings including Python, Node.js, Go, and Swift, or use it as a CLI, self-hosted REST API, or MCP server. Docker deployment is supported, and its OCR backends cover Linux, macOS, Windows, and the browser through WebAssembly.

For scanned documents, xberg supports Tesseract, PaddleOCR, and Candle, with fallback between backends and recognition confidence scores. Layout and table models reconstruct reading order and cell structure for Markdown output. Whisper ONNX handles speech transcription from audio files and video audio tracks. It also extracts functions, imports, and docstrings from source code, with chunks split along syntax boundaries for AI retrieval.

Structured extraction produces schema-validated JSON using Ollama, LM Studio, or vLLM locally, or hosted models from providers such as OpenAI and Anthropic. Embeddings and reranking also have local ONNX and hosted options; choosing a hosted provider sends that work to an external service. Other capabilities include personal-data redaction, keyword extraction, and LLM-based summaries and translation. Output includes plain text, Markdown, HTML, JSON trees, and Docling DocTags.

Similar to xberg (formerly Kreuzberg)