30.1KUpdated 3 weeks agoMIT
macOS#GGUF#Hugging Face integration#Hybrid search
QMD is a local search engine for people with Markdown notes, meeting transcripts, or documentation they want to search themselves or make available to an AI agent. It accepts exact keywords and natural-language queries, with indexing and model inference running on your own machine. It's open source under MIT.
12KUpdated 1 week agoApache-2.0
Docker · Web#Batch processing#Human approval#Multi-agent workflows
Bisheng is an open source, self-hosted platform for teams building AI applications around business documents and processes. Its visual workflow editor combines automated tasks with human feedback, including intervention during multi-turn conversations. It's suited to document review, support ticket assistance and report generation that need more control than a single chatbot exchange.
perfectmemory.aiAI Notes and Knowledge Bases
macOS · Windows#LM Studio integration#MCP#Ollama integration
Perfect Memory AI keeps a searchable record of what you've seen on your screen and heard in meetings on Mac and Windows. It's for people who need to find a page, presentation or conversation again without remembering which app held it. Recordings stay on your device, and search returns material you've actually seen or heard.
3.2KUpdated 3 months agoMIT
#Guardrails
LLM Guard is a Python security toolkit for developers building applications around large language models. It checks prompts and generated responses for risks such as prompt injection, sensitive data exposure and harmful language. The project is archived and no longer maintained, including its associated models on Hugging Face.
1.5KUpdated 3 years agoApache-2.0
Web#Guardrails#Semantic search
Rebuff is a prompt injection detector for developers building LLM applications that accept untrusted input. It combines checks for suspicious prompts with a record of past attacks and tests for leaked prompt content. The project is archived and no longer maintained.
496Updated 3 years agoApache-2.0
Docker · Web#Guardrails#Semantic search
Vigil is a self-hosted security scanner for developers and researchers who want to check LLM inputs and responses for prompt injection, jailbreak attempts, and other suspicious content. It combines several detection methods and includes attack signatures and datasets, so teams can assess known threats without building every detector themselves. It is experimental alpha software for research and is open source under Apache 2.0.
5.6KUpdated 2 years agoApache-2.0
#Ollama integration#OpenAI-compatible API
RouteLLM is a self-hosted Python framework for developers who want to split requests between a stronger LLM and a cheaper model. It judges which prompts need the stronger model, so an application doesn't have to send every request to its most expensive provider. You control the cost-quality tradeoff through a routing threshold.
318Updated 3 weeks agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Quantization
picoLLM is an on-device inference SDK for developers building apps that run compressed language models on users' hardware. It generates text locally, so prompts don't need to go to a cloud inference service. Its main distinction is Picovoice's compression method, which learns how to allocate precision across model weights rather than applying a fixed allocation.
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.
30.4KUpdated 18 hours agoMIT
#MCP#Tool calling
Composio connects AI assistants and custom agents to apps such as Gmail, Slack, GitHub, and Linear. It's for people who want their assistant to act on requests across apps, and developers who don't want to maintain each integration themselves. Its CLI gives coding agents a local interface; the standard setup uses Composio's hosted authentication and execution service. It requires an account and internet access.
3.8KUpdated 2 years agoMIT
Docker#Streaming inference#Tool calling
Vocode is an open source Python library for developers building voice AI agents, with a self-hosted telephony server and support for live conversations through a computer's microphone and speakers. It connects speech recognition, an LLM, and speech synthesis in one library. The code uses the MIT license.
90.7KUpdated 1 week ago
Windows#Git integration#Knowledge graphs#MCP
MCP Reference Servers is a collection of locally run examples for developers building connections between AI applications and external tools or data. Maintained by the MCP steering group, the servers demonstrate the protocol and its SDKs. They're educational implementations, so developers should assess security requirements before using them in production.
2.9KUpdated 1 day agoMIT
Docker#MCP
Supergateway connects MCP servers and AI clients that use different connection types. It's for developers who want to expose a local tool server over a network or connect a desktop client to a remote server. The gateway runs on your own machine or server through Node.js or Docker, and its code is open source under the MIT license.
2.8KUpdated 5 months agoMIT
macOS · Windows · Linux · Docker#MCP
mcp-proxy is a self-hosted tool for connecting AI clients and MCP servers that use different connection types. It's for developers who need a stdio client, such as Claude Desktop, to reach an SSE or Streamable HTTP server, or who want remote clients to access a local stdio server. The proxy runs on your own machine or server.
23.8KUpdated 8 months agoMIT
Web#LLM tracing#Multi-user access#Ollama integration
Vanna is a self-hosted Python framework for building AI agents that answer database questions in plain language. It's for teams adding chat to analytics products or internal data tools, especially when each user needs different access to the data. The project is archived and no longer maintained. Its open-source code uses the MIT license.
10.3KUpdated 1 year agoMIT
macOS · Windows · Linux#Multimodal input#Ollama integration
Self-Operating Computer lets a vision-capable AI model control a desktop by reading the screen and choosing mouse and keyboard actions to carry out a goal. It's a Python framework for developers and researchers exploring AI agents that work through application interfaces. The code uses the MIT license.
6.4KUpdated 2 years agoApache-2.0
Web · Browser Extension#Code execution
LaVague is a Python framework for developers building AI agents that carry out tasks in a web browser. An agent takes a plain-language objective, examines the current page, and generates and executes browser actions across multiple steps. It's open source under the Apache 2.0 license.
getzep/zepAgent Memory
Docker#Hybrid search#Knowledge graphs#OpenAI-compatible API
Zep Community Edition v1.0.2 is a deprecated, unsupported self-hosted memory service for developers building AI agents and conversational assistants. This legacy edition uses the Apache 2.0 license. It turns chat history into a knowledge graph that records how facts change over time, so an assistant can distinguish a user's current preferences from earlier ones.
6.9KUpdated 1 year agoApache-2.0
Web#Multi-agent workflows#Multilingual#RAG
MindSearch is a self-hosted AI search framework for people who want to build their own Perplexity-style answer engine. Multiple LLM agents search and read web pages to produce answers with a visible research process. It's open source under Apache 2.0.
4.7KUpdated 1 month agoMIT
macOS · Windows · Linux#Visual workflows
Rivet is a desktop visual programming environment for developers building AI agents and applications with complex LLM workflows. Its editor runs on macOS, Windows and Linux, and its TypeScript library executes the resulting graphs inside your own application. The project uses the MIT license.
55.5KUpdated 2 months ago
Docker · Web#Human approval#LLM tracing#Multi-agent workflows
Flowise is a visual builder for AI agents and chatbots that can run locally or on your own server, including through Docker. It's for developers and teams building LLM applications with connected workflow blocks. The project is archived and no longer maintained.
hillnote.comAI Notes and Knowledge Bases
macOS · Windows · iOS · Android#MCP#Ollama integration#Works offline
Hillnote is a writing and planning workspace for people who want their notes on their own disk, with AI available inside the editor. It runs on Mac, Windows, iOS and Android. The editor, workspace and local AI work offline, and getting started doesn't require an account.
35Updated 7 months agoMIT
macOS · Web · VS Code#Code execution#LM Studio integration#MCP
Wingman-AI is a self-hosted AI agent platform for work that needs ongoing context and several agents with different roles. It suits developers and teams handling research, support, or recurring operations alongside coding tasks. A lead agent can delegate work to specialized subagents, each with its own workspace and session history.
4.2KUpdated 1 year agoApache-2.0
Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
pieces.appAgent Memory
#MCP#Persistent memory#RAG
Pieces runs in the background on your computer and builds a searchable history of your work across apps. It's for people who need to recover a decision, find earlier research or resume a task after switching between tools. Memories stay on your machine by default.
5.8KUpdated 4 months agoPostgreSQL
Docker#Batch processing#Ollama integration#RAG
pgai keeps search embeddings in sync with PostgreSQL data for developers building RAG applications and AI agents. It's a Python library with database components and workers you can self-host, including in Docker. The project is archived and no longer maintained or supported. Its code is open source under the PostgreSQL License.
1.1KUpdated 3 weeks agoMPL-2.0
Linux#Ollama integration#RAG#Semantic search
chromem-go is a vector database that runs inside your Go application, so developers can add semantic search or retrieval augmented generation (RAG) without maintaining a separate database server. It stores text alongside embeddings and retrieves related documents for use in LLM answers. Its focus is ordinary application workloads rather than collections containing millions of documents.
39.6KUpdated 1 year ago
#Ollama integration#RAG#Reranking
Quivr Core is a Python framework for developers adding document-based AI answers to their own applications. It combines file ingestion with retrieval-augmented generation (RAG), so a model can answer questions using material from your documents. It supports local models through Ollama as well as cloud APIs from OpenAI, Anthropic, and Mistral.
187.6KUpdated 3 days ago
Docker · Web#Human approval#Multi-agent workflows#Scheduled tasks
AutoGPT is an AI agent platform for teams handling recurring business work, with a hosted service and a self-hosted option. Its coordinator, Otto, assigns work to specialists for content, sales, research and operations. You can describe an outcome in plain English or talk directly to the specialist responsible for it.
61.2KUpdated 6 months agoCC-BY-4.0
Web#Code execution#MCP#Multi-agent workflows
AutoGen is a framework for developers building AI agents that work together or alongside people. Its agent runtime can run locally or across distributed systems, while model integrations such as OpenAI and Azure OpenAI send requests to external services. It's community-managed and in maintenance mode, with no further features or enhancements planned.
17.7KUpdated 2 years agoMIT
Docker · Web#Human approval#Persistent memory#Tool calling
SuperAGI is a self-hosted AI agent framework for developers who want to build task automation and manage agents through a browser interface. It runs locally in Docker and can use local LLMs with a GPU. The Python project uses the MIT license, so you can modify it for your own applications.
8.4KUpdated 21 hours agoMIT
#Agent Skills#Batch processing#Guardrails
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
20.1KUpdated 2 days agoMIT
macOS · Linux · Docker · Web#Code execution#llama.cpp backend#OpenAI-compatible API
DB-GPT is a self-hosted AI data assistant for teams analyzing business data and developers building data applications. It turns plain-language requests into SQL queries and Python analysis, then produces charts, dashboards, or HTML reports. You can run it on macOS or Linux, with Docker deployment also supported.
3.1KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows
BotSharp is a self-hosted framework for .NET developers building AI agents into business applications. Written in C#, it runs on Windows, Linux and macOS and is open source software under Apache 2.0. Its plugin design lets teams choose their model provider, storage and interface while keeping agent coordination in the same framework.
9.7KUpdated 9 months agoMIT
#Ollama integration#RAG#Semantic search
LangChainGo is a Go implementation of LangChain for developers building LLM applications in their own software. It connects Go programs to model backends, including Ollama for local LLM use and cloud services such as OpenAI and Gemini. It's a library, so its audience is developers who want to build an application rather than use a ready-made chat interface.