
Gitingest turns a Git repository or local directory into a text digest that developers can give to an LLM as code context. It combines the directory tree and file contents in one extract, so you don't have to assemble context file by file. It's open source under the MIT license.
You can run it locally through a command-line tool or use it as a Python package in your own code. The web interface is available as a hosted service, and you can self-host it with Docker. Local directory processing lets you prepare code on your machine; the hosted website clones and processes repositories on its backend.
The digest includes a summary, directory structure and file contents, with statistics for extract size and estimated token count. These help you judge how much code you're about to put into a prompt. Include and exclude filters let you choose the relevant files, and the command-line tool skips files covered by .gitignore by default. The website lets you copy individual sections or download the digest.
The output is plain text for use with any LLM, so Gitingest doesn't tie the extract to a particular model provider. Its Python package also lets developers make codebase extraction part of a larger application or workflow.
Private GitHub repositories require a personal access token. On the hosted service, the backend uses that token to clone the repository, then discards it from memory without storing it. The website doesn't cache the token in the browser, and it deletes cloned repositories after processing.
Claim this page with an email at gitingest.com. Gitingest gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Gitingest?Promote it
Something wrong or outdated on this page?
12KUpdated 1 week agoApache-2.0
Docker · Web#Batch processing#Human approval#Multi-agent workflows
Bisheng is an open source, self-hosted platform for teams building AI applications around business documents and processes. Its visual workflow editor combines automated tasks with human feedback, including intervention during multi-turn conversations. It's suited to document review, support ticket assistance and report generation that need more control than a single chatbot exchange.
6.4KUpdated 1 day agoApache-2.0
Docker · Web
docTR is an open-source Python OCR library for developers building document processing tools and researchers comparing text recognition models. It reads PDFs and images on your own hardware, locating words and recognizing their text. The library uses PyTorch and carries the Apache 2.0 license.
39.9KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#Knowledge graphs#LLM tracing#Multimodal input
10.9KUpdated 3 years agoMIT
macOS · Docker · Web#GGUF#llama.cpp backend#OpenAI-compatible API
6.7KUpdated 2 months agoApache-2.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
3.7KUpdated 5 days ago
Docker · Web#MCP#Multi-user access#Multimodal input
Morphik Core is a self-hosted multimodal retrieval engine for developers building AI applications around visually rich documents. It searches diagrams, schematics, charts, and datasheets alongside text, so applications can retrieve information that text extraction alone can miss. You can run it on your own server, including through Docker, or use Morphik's hosted service.
LightRAG combines knowledge graphs with vector search to answer questions across a document collection. It's a self-hosted Python framework for developers building document assistants, particularly where answers depend on relationships between facts in different files, such as legal or financial material.
LlamaGPT is a self-hosted ChatGPT alternative for people who want general chat or coding help on their own computer or home server. It runs models locally and keeps conversation data on your device. After the initial model download, it works offline.
MonkeyOCR is a local AI document parser for developers and researchers working with English and Chinese PDFs or images. It extracts text, formulas and tables while identifying page structure and relationships between blocks. That makes it useful for documents where plain text extraction loses reading order or separates content from its layout.