15.9KUpdated 7 months agoApache-2.0
Ragas is an open-source Python library for developers who need repeatable evaluations of LLM applications and retrieval-augmented generation (RAG) systems. It combines model-based scoring with traditional metrics so teams can compare application changes using test results rather than manual judgments alone. Its license is Apache 2.0.
12.5KUpdated 1 month agoApache-2.0
Web#LLM tracing#Multi-user access#Streaming inference
Chainlit is an open-source Python framework for developers building conversational AI apps with their own application logic. You can run the app on your own server and give users a browser chat interface. It's a framework for creating an app, rather than a ready-made chatbot with a fixed model.
31.2KUpdated 22 hours agoApache-2.0
Docker · Web#Knowledge graphs#MCP#Multi-user access
Cognee gives AI agents persistent memory across sessions, connecting documents, code, and conversations in a searchable knowledge graph. It's for developers who want agents to retain project context and teams whose knowledge sits across tickets, discussions, and repositories. The Python package is open source under Apache 2.0.
13.2KUpdated 1 day agoApache-2.0
#MCP#Ollama integration#RAG
LangChain4j is an Apache 2.0 open-source Java library for developers building chatbots, assistants and AI agents in JVM applications. It connects application code to local LLM backends such as Ollama as well as cloud providers such as OpenAI and Google Vertex AI. Where model requests go depends on the backend you choose.
4.4KUpdated 1 day agoMIT
#Human approval#Multi-agent workflows#Multimodal input
RubyLLM is an MIT-licensed AI framework for developers building Ruby and Rails applications with local or hosted models. Its shared API lets an application switch between Ollama, cloud providers such as Anthropic and OpenAI, and OpenAI-compatible endpoints without rewriting its model integration. The framework runs in your application; model processing happens at the local or hosted backend you choose.
28.6KUpdated 2 days agoMIT
Docker · Web · Browser Extension#Agent Skills#Code execution#MCP
Repomix turns a code repository into a single file that an AI assistant can read. It’s for developers who need to give Claude, ChatGPT, Gemini, or another model enough project context for code review or analysis. You can use its command line on your own machine, pack a repository through its website, or run it as an MCP server, including through Docker. It is open source under the MIT license.
2.8KUpdated 2 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#Git integration#Human approval
Vexa is a meeting bot and transcription API for developers building meeting features and teams feeding calls into AI agents. Bots join Google Meet, Microsoft Teams and Zoom, then send live transcripts with speaker labels to your application or agent. You can host the full platform on your own infrastructure or use Vexa's cloud service.
1.4KUpdated 2 months agoMIT
#Multimodal input#Ollama integration#Streaming inference
OllamaSharp is a C# library for developers building .NET applications around Ollama. It connects to Ollama on your own machine or a remote server and covers the full Ollama API, including model management alongside chat and embeddings. It's open source under the MIT license.
14KUpdated 3 weeks agoMIT
#llama.cpp backend#Ollama integration#Streaming inference
Instructor is an open-source library for developers who need structured data from local LLMs or cloud models. It turns natural-language input into typed objects that applications can use, with validation and retries built into the extraction process. Its focus is data extraction.
22.9KUpdated 3 days agoGPL-3.0
Docker · Web#MCP#Multimodal input#RAG
MaxKB is a self-hosted AI agent platform for organizations building customer support bots, internal knowledge assistants and business automation. It combines answers grounded in company documents with workflows that can call functions and MCP tools. You can run it on your own server through Docker and use it through a browser.
15.9KUpdated 1 month agoApache-2.0
#llama.cpp backend#MLX#Ollama integration
Outlines is an open-source Python library for developers who need LLM responses to match a defined structure. It constrains output during generation, reducing the need to repair malformed JSON or retry responses that don't fit an application's requirements. The library uses the Apache 2.0 license.
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.
84.5KUpdated 1 day agoBSD-3-Clause
Docker#Guardrails#MCP#Tool calling
Scrapling is a Python web scraping framework for developers collecting website data or giving AI agents access to web pages. Its adaptive parser can find previously selected elements after a site's layout changes, reducing the need to repair extraction rules. It's open source under the BSD-3-Clause license and runs on your own machine or in Docker.
1.3KUpdated 1 month ago
Web#Human approval#MCP#Multi-user access
Hexabot is a self-hosted AI workflow automation platform for teams building customer service assistants and business automations. It connects conversations to actions across websites, messaging platforms, social channels and custom entry points. You can run it on-premise or in your own cloud, with workflows, memory and customer data in infrastructure you control.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
29.8KUpdated 24 hours agoMIT
macOS · Windows · Linux#Code execution#Guardrails#Human approval
OpenAI Agents SDK is an open-source Python framework for developers building AI apps that need to use tools, delegate tasks, or work across multiple steps. Its runtime manages agent turns and conversation state while letting developers express workflows in ordinary Python. It uses the MIT license.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
3.8KUpdated 24 hours agoApache-2.0
#Hugging Face integration#Quantization
LLM Compressor is an open-source Python library for developers preparing models to run on their own hardware with vLLM. It reduces model size and memory requirements through quantization, and accepts local checkpoints or models from Hugging Face repositories. It's licensed under Apache 2.0.
19.2KUpdated 2 weeks agoApache-2.0
Web · Browser Extension#OpenAI-compatible API#Streaming inference#Structured output
WebLLM runs language models directly in a user's browser, using WebGPU for GPU acceleration. It's an open-source engine for developers building web-based AI assistants and Chrome extensions that process prompts on the user's device rather than an inference server. The project uses the Apache 2.0 license.
82.9KUpdated 3 hours ago
Linux · Docker · Web#MCP#Multi-agent workflows#Multi-user access
LobeHub is an AI agent workspace for people who want to assign work to several assistants and keep that work organized. It offers a hosted service and a Docker-based self-hosted version for a private device or server. A Linux download is also available.
26.6KUpdated 1 day agoApache-2.0
Docker#Guardrails#Hugging Face integration#Hybrid search
Haystack is a Python framework for developers building self-hosted AI agents, document search, and apps that answer questions using their own data. Its modular pipelines let teams control which information reaches a model and inspect how retrieval, memory, tools, and generation contribute to an answer. It's open source under Apache 2.0.
7.2KUpdated 23 hours ago
Docker#Guardrails#LLM tracing#Tool calling
NeMo Guardrails is an open-source Python toolkit for developers who need control over how an AI assistant responds and uses tools. It runs within your application or as a self-hosted server, including in Docker. The library uses the Apache 2.0 license.
10.6KUpdated 2 days agoMIT
#Batch processing#Ollama integration#Streaming inference
Ollama Python connects Python applications to models running through Ollama on your own machine or to Ollama's cloud service. It's for developers adding local LLM features to scripts, chat applications, or other Python projects. The library is open source under the MIT license and requires a running Ollama service for local use.
7.7KUpdated 2 days agoApache-2.0
macOS · Docker · Web
Steel is a browser API for developers building AI agents that interact with websites. It runs Chrome sessions locally or on a self-hosted server through Docker, so you can keep browser infrastructure on hardware you control. The code uses the Apache 2.0 license. Steel also offers a hosted service where browser sessions run in its cloud.
7.4KUpdated 3 weeks agoLGPL-3.0
#Hugging Face integration#LoRA
mergekit combines existing language models into a single model on your own hardware, without additional training or access to the original training data. It's for developers and researchers who want to combine fine-tuned capabilities or adjust the balance between model behaviors. The Python toolkit is open source under LGPL-3.0.
27.1KUpdated 2 hours ago
#Agent Skills#Streaming inference#Structured output
Vercel AI SDK is a TypeScript library for developers building chatbots, generative interfaces, and AI agents. It gives an application one way to work with models from providers such as OpenAI, Anthropic, and Google, so teams can change providers without rebuilding the parts of their app that handle responses. It suits developers who want control over the application they build while choosing how it connects to model services.
43.6KUpdated 1 day agoApache-2.0
Web#MCP
Gradio turns a Python function, model, or API into a web interface. It runs on your computer, making it useful for developers and researchers who want people to try an AI project through a browser. It's open source under the Apache 2.0 license.
470Updated 1 month agoMIT
macOS#MCP#Tool calling
ast-grep MCP is a locally run server that lets AI coding assistants search code by its syntax structure. It connects ast-grep to Cursor, Claude Desktop and other clients that support the Model Context Protocol (MCP). It's for developers who need an assistant to find specific code constructs when matching words alone is too broad.
42Updated 13 hours ago
macOS · Windows · Linux · Docker#Agent Skills#Code execution#Distributed execution
Code Buddy is an AI coding assistant for developers who want a terminal agent running on their own machine, with a choice of local or cloud models. It reads repositories, edits code and runs commands on Linux, macOS and Windows. Ollama keeps model inference local without an API key or account; cloud providers send model requests off the machine. Routing includes automatic provider failover.