Quivr Core is a Python framework for developers adding document-based AI answers to their own applications. It combines file ingestion with retrieval-augmented generation (RAG), so a model can answer questions using material from your documents. It supports local models through Ollama as well as cloud APIs from OpenAI, Anthropic, and Mistral.
The framework provides a starting workflow while leaving retrieval behavior open to customization. You can adapt how it handles conversation history, rewrites questions, retrieves relevant passages, and generates answers. It also supports reranking through Cohere and streaming answers to users.
Quivr Core accepts PDF, TXT, and Markdown files, with custom parsers for other formats. Its Megaparse integration provides another route for document ingestion. PGVector and Faiss are supported vector stores, giving developers a choice of where to keep the searchable document index.
Model choice affects where processing happens: Ollama runs models locally, while the supported cloud APIs send model requests to external providers. Internet search and tool use can extend the document workflow when an application needs information or actions beyond its files.
The Apache 2.0 licensed core library can be embedded in an existing Python application. Developers can change retrieval strategies independently of the application's user interface.
Claim this page and we'll verify you by hand. Quivr Core gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Quivr Core?Promote it
Something wrong or outdated on this page?
26.6KUpdated 1 day agoApache-2.0
Docker#Guardrails#Hugging Face integration#Hybrid search
Haystack is a Python framework for developers building self-hosted AI agents, document search, and apps that answer questions using their own data. Its modular pipelines let teams control which information reaches a model and inspect how retrieval, memory, tools, and generation contribute to an answer. It's open source under Apache 2.0.
13.2KUpdated 1 day agoApache-2.0
#MCP#Ollama integration#RAG
9.7KUpdated 9 months agoMIT
#Ollama integration#RAG#Semantic search
52.4KUpdated 2 days agoMIT
#Ollama integration#RAG#Reranking
5.8KUpdated 4 months agoPostgreSQL
Docker#Batch processing#Ollama integration#RAG
4.4KUpdated 1 day agoMIT
#Human approval#Multi-agent workflows#Multimodal input
LangChain4j is an Apache 2.0 open-source Java library for developers building chatbots, assistants and AI agents in JVM applications. It connects application code to local LLM backends such as Ollama as well as cloud providers such as OpenAI and Google Vertex AI. Where model requests go depends on the backend you choose.
LangChainGo is a Go implementation of LangChain for developers building LLM applications in their own software. It connects Go programs to model backends, including Ollama for local LLM use and cloud services such as OpenAI and Gemini. It's a library, so its audience is developers who want to build an application rather than use a ready-made chat interface.
LlamaIndex is an MIT-licensed Python framework for developers building AI agents and apps that answer questions using their own data. It connects documents and other sources to language models, then helps an app find the relevant material when a user asks something. Its open source framework can work with models served through Ollama.
pgai keeps search embeddings in sync with PostgreSQL data for developers building RAG applications and AI agents. It's a Python library with database components and workers you can self-host, including in Docker. The project is archived and no longer maintained or supported. Its code is open source under the PostgreSQL License.
RubyLLM is an MIT-licensed AI framework for developers building Ruby and Rails applications with local or hosted models. Its shared API lets an application switch between Ollama, cloud providers such as Anthropic and OpenAI, and OpenAI-compatible endpoints without rewriting its model integration. The framework runs in your application; model processing happens at the local or hosted backend you choose.