52.4KUpdated 2 days agoMIT
#Ollama integration#RAG#Reranking
LlamaIndex is an MIT-licensed Python framework for developers building AI agents and apps that answer questions using their own data. It connects documents and other sources to language models, then helps an app find the relevant material when a user asks something. Its open source framework can work with models served through Ollama.
29.8KUpdated 1 day ago
Docker · Web#LLM tracing#MCP#Multi-user access
FastGPT is a self-hosted AI agent builder for teams that want assistants to answer questions using company documents and carry out business workflows. Its visual editor connects model calls, knowledge retrieval and tools into applications for customer support, internal knowledge search and document review. You can run the platform on your own server through Docker or use the vendor's hosted service.
16.9KUpdated 1 day ago
Docker#Hybrid search#RAG#Reranking
Weaviate is a self-hosted vector database for developers building search applications, RAG systems, recommendation engines, and chatbots. It stores data objects alongside their vector embeddings, so applications can search by meaning and filter results using structured data. You can run the database locally with Docker, deploy it on Kubernetes, or use the hosted Weaviate Cloud service.
31.2KUpdated 22 hours agoApache-2.0
Docker · Web#Knowledge graphs#MCP#Multi-user access
Cognee gives AI agents persistent memory across sessions, connecting documents, code, and conversations in a searchable knowledge graph. It's for developers who want agents to retain project context and teams whose knowledge sits across tickets, discussions, and repositories. The Python package is open source under Apache 2.0.
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
13.2KUpdated 1 day agoApache-2.0
#MCP#Ollama integration#RAG
LangChain4j is an Apache 2.0 open-source Java library for developers building chatbots, assistants and AI agents in JVM applications. It connects application code to local LLM backends such as Ollama as well as cloud providers such as OpenAI and Google Vertex AI. Where model requests go depends on the backend you choose.
4.4KUpdated 1 day agoMIT
#Human approval#Multi-agent workflows#Multimodal input
RubyLLM is an MIT-licensed AI framework for developers building Ruby and Rails applications with local or hosted models. Its shared API lets an application switch between Ollama, cloud providers such as Anthropic and OpenAI, and OpenAI-compatible endpoints without rewriting its model integration. The framework runs in your application; model processing happens at the local or hosted backend you choose.
19.7KUpdated 6 months agoApache-2.0
Linux · Docker · Web#Batch processing#Distributed execution#OpenAI-compatible API
olmOCR is an open-source OCR toolkit for turning PDFs and image documents into text for LLM datasets and training. It suits researchers and developers who need readable document content, including pages where columns, figures, or mathematical notation make text extraction difficult. You can run it on your own GPU, including through Docker, or use a remote inference server.
22.9KUpdated 3 days agoGPL-3.0
Docker · Web#MCP#Multimodal input#RAG
MaxKB is a self-hosted AI agent platform for organizations building customer support bots, internal knowledge assistants and business automation. It combines answers grounded in company documents with workflows that can call functions and MCP tools. You can run it on your own server through Docker and use it through a browser.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
9.3KUpdated 2 months agoApache-2.0
#Multilingual#Multimodal input#RAG
PaperQA2 is an open source Python research assistant for people who need answers grounded in a collection of scientific papers. It searches documents on your machine and writes answers with in-text citations, including page references. Researchers can use it to summarize findings or check for contradictions across papers, while developers can build it into their own research tools.
46.3KUpdated 1 day agoApache-2.0
macOS · Linux#Hybrid search#Semantic search
Milvus is an open-source vector database for developers building RAG applications, image search and recommendation systems. It stores embeddings alongside metadata so applications can retrieve related text, images or multimodal data. You can run it on your own hardware, from a laptop prototype to a distributed production cluster.
9.3KUpdated 21 hours agoAGPL-3.0
Docker#MCP#Semantic search
ParadeDB adds full-text, vector and hybrid search to the Postgres database that holds your application data. It's for developers who need search and analytics without maintaining a separate Elasticsearch cluster or synchronizing a second copy of their data. It runs as the pg_search extension in self-managed Postgres, including local Docker deployments.
5.1KUpdated 22 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
23.2KUpdated 1 day ago
macOS · Windows · Linux · Docker#Semantic search
pgvector adds vector storage and similarity search to Postgres, so developers can keep embeddings alongside application records in a self-hosted database. It suits applications that need to find similar items while retaining SQL queries, joins and transactional guarantees. It runs on Linux, macOS and Windows, with Docker also supported.
3.2KUpdated 24 hours agoApache-2.0
#Batch processing#Multilingual#ONNX
FastEmbed is a Python library that generates embeddings on your own hardware for semantic search and retrieval-augmented generation (RAG). It's for developers who need to turn text into searchable vectors without relying on a cloud embedding API. It can run on a CPU or use GPU acceleration, and its Apache 2.0 license makes it open source.
21.4KUpdated 3 weeks agoApache-2.0
macOS · Web#Batch processing#llama.cpp backend#Multilingual
Surya is a local OCR toolkit for developers extracting text and structure from PDFs and document images. It combines text recognition, layout analysis and table recognition in one vision-language model, so results retain page structure and reading order rather than just the words.
90.4KUpdated 2 weeks agoApache-2.0
Web#Multilingual#ONNX#Structured output
PaddleOCR is an open source OCR and document parsing toolkit for developers building document search, RAG systems and AI agents. It runs on your own hardware or a self-hosted server and turns PDFs and images into structured Markdown or JSON. The Python toolkit uses PaddlePaddle and carries the Apache 2.0 license.
7.1KUpdated 1 day agoApache-2.0
Linux#Hybrid search#RAG#Semantic search
Vespa is a self-hosted AI search platform for developers building search, RAG, and recommendation systems over large, changing datasets. It combines retrieval with machine-learned ranking, so an application can find candidate results and evaluate their relevance in the same platform. The code is open source under Apache 2.0. You can run it on your own servers or use the managed Vespa Cloud service, where applications run in the cloud.
26.6KUpdated 1 day agoApache-2.0
Docker#Guardrails#Hugging Face integration#Hybrid search
Haystack is a Python framework for developers building self-hosted AI agents, document search, and apps that answer questions using their own data. Its modular pipelines let teams control which information reaches a model and inspect how retrieval, memory, tools, and generation contribute to an answer. It's open source under Apache 2.0.
19.1KUpdated 1 week agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
Sentence Transformers is an open-source Python library for developers building semantic search and document retrieval on their own hardware. It runs embedding and reranker models locally, turning content into numerical representations for similarity comparisons and scoring results against a query. The library uses the Apache 2.0 license.
2.9KUpdated 6 months agoMIT
macOS · Docker#Batch processing#Hugging Face integration#Multimodal input
Infinity Embeddings is a self-hosted server for developers building semantic search and retrieval-augmented generation applications. It runs embedding and reranking models on your own hardware, with support for image and audio search alongside text. It's open source under MIT.
37.5KUpdated 2 months agoAGPL-3.0
Web#RAG#Semantic search#Web search
Khoj is an AI assistant for people who want to ask questions across their own files, research the web, and give recurring work to agents. You can self-host it on your computer or server, or use Khoj's cloud app. It's open source under the GNU AGPL v3.0 license.