
Chonkie is an MIT-licensed, open-source library for developers preparing documents for retrieval-augmented generation (RAG). It combines text cleaning, chunking, embeddings and vector database ingestion, so you don't have to assemble each stage from separate libraries. It runs locally or in the cloud, with Python and JavaScript support and a self-hosted REST API that can run in Docker.
Its chunking methods include fixed-size token splits, sentence-based splits and semantic chunking. A byte-based chunker uses SIMD acceleration for fast text splitting. The library keeps its dependency footprint small by making specialized components optional, and it supports multilingual text.
Document processing covers local files and directories, structured Markdown, and CSV or Excel tables. After splitting, Chonkie can merge overlapping chunks based on similarity and attach embeddings through a chosen provider. Reusable pipelines connect these stages, while vector database integrations handle ingestion. You can also export chunks as JSON or HuggingFace datasets.
Local processing and the self-hosted API run on your own hardware. Some integrations use external services: Mistral OCR sends images or PDFs to the Mistral OCR API for text extraction. Chonkie also provides skills for coding assistants including Claude Code, Cursor and Copilot, covering its APIs and RAG pipeline patterns.
Claim this page with an email at docs.chonkie.ai. Chonkie gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Chonkie?Promote it
Something wrong or outdated on this page?
52.4KUpdated 14 hours agoMIT
#Ollama integration#RAG#Reranking
LlamaIndex is an MIT-licensed Python framework for developers building AI agents and apps that answer questions using their own data. It connects documents and other sources to language models, then helps an app find the relevant material when a user asks something. Its open source framework can work with models served through Ollama.
38.5KUpdated 17 hours agoMIT
#RAG#Source citations#Streaming inference
13KUpdated 18 hours agoApache-2.0
Docker#Agent Skills#Hugging Face integration#Knowledge graphs
15.5KUpdated 5 hours agoApache-2.0
macOS · Windows · Linux · Docker#Multilingual
12KUpdated 1 week agoApache-2.0
Docker · Web#Batch processing#Human approval#Multi-agent workflows
1.6KUpdated 2 weeks agoApache-2.0
macOS · Windows · Linux · Docker#MCP#Multimodal input#Ollama integration
PageIndex is an open-source document RAG engine for developers building question answering over long PDFs, such as financial reports, legal documents and technical manuals. It organizes documents into a hierarchical tree and uses an LLM to find relevant sections, rather than relying on vector similarity search. It doesn't require a vector database or document chunking.
txtai is a Python framework for developers building search applications, chat with their data, and AI agents on their own hardware or servers. Its embeddings database combines sparse and dense vector search with graphs and relational data, so the same system can find related content and supply context to language models. It's open source under Apache 2.0.
Unstructured is a local document processing library for developers building LLM applications and document ingestion pipelines. It turns PDFs, Word documents, HTML, emails and images into document elements that applications can use. The Python library is open source under Apache 2.0 and runs on your own hardware, including through Docker images for x86_64 and Apple Silicon.
Bisheng is an open source, self-hosted platform for teams building AI applications around business documents and processes. Its visual workflow editor combines automated tasks with human feedback, including intervention during multi-turn conversations. It's suited to document review, support ticket assistance and report generation that need more control than a single chatbot exchange.
ExtractThinker is a Python library for developers who need structured data from documents, such as invoice fields their application can use. It pairs document parsers with an LLM and returns results that follow a Pydantic schema. It's open source under Apache 2.0 and runs on macOS, Windows and Linux.