7.4KUpdated 12 hours agoApache-2.0
Linux#Agent Skills#OpenAI-compatible API#Prompt versioning
Reef is self-hosted infrastructure for developers who want AI agents to improve through feedback on actual interactions. It connects inference and learning with versioned deployment, so an agent can update its model weights or its prompts, rules, and skills while continuing to serve requests. It's open source under Apache 2.0.
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
756Updated 5 hours agoMIT
macOS · Windows · Linux#LLM tracing#MCP
agentacct is a local dashboard for developers who want to review what their coding agents did across projects and compare usage with recorded results. It brings task activity, file changes and test outcomes together with token counts and estimated costs. Claude Code and Codex have the fullest task records; it also supports OpenCode, Kimi Code and Hermes, with coverage varying by capability.
29.8KUpdated 1 day agoApache-2.0
macOS · Docker · Web#Code execution#LLM tracing#Multi-user access
Sim is an AI agent workspace for teams that need to build automations and control how people use them across an organization. You can self-host it on your own server or cloud with Docker or Kubernetes. Its open-source code uses the Apache 2.0 license.
28.4KUpdated 21 hours ago
Web#Human approval#LLM tracing#MCP
Mastra is a TypeScript framework for developers building AI agents and applications on their own servers or inside existing web apps. Its server runs locally or as a standalone deployment; Mastra Cloud provides a hosted alternative. Model routing connects to providers such as OpenAI, Anthropic, and Gemini, so running the framework locally doesn't keep those model requests on your machine.
157.6KUpdated 3 hours ago
Docker · Web#Code execution#MCP#OpenAI-compatible API
Dify is a source-available platform for teams building AI agents and apps on a visual canvas. Its Community Edition runs on your own server with Docker. Dify also offers a hosted cloud service, while Enterprise deployments can run in a VPC or on a self-hosted server. The Community Edition uses a custom Apache 2.0 derivative license.
22.3KUpdated 1 day agoApache-2.0
Web#Guardrails#LLM tracing#Prompt versioning
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
27.3KUpdated 20 hours agoMIT
macOS · Windows · Linux#Code execution#MCP#Tool calling
Cua gives AI agents access to computers they can inspect and operate, with tools for desktop automation, local virtual machines, and hosted fleets. It's for developers building agents that work across native apps and browsers, or evaluating how well those agents complete computer tasks. You bring the agent and model.
35.2KUpdated 2 hours ago
Docker#Agent Skills#LLM tracing#MCP
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.
20.3KUpdated 1 hour agoMIT
#Human approval#LLM tracing#MCP
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
59.9KUpdated 57 minutes ago
#Guardrails#MCP#Multi-user access
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.
25.6KUpdated 1 hour agoMIT
#AI red teaming#Git integration#MCP
Promptfoo is an open source CLI and library for testing prompts, AI agents, and RAG applications. It runs evaluations locally and helps developers compare model responses while security teams look for weaknesses in the applications built around them. The project is MIT licensed.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
3.3KUpdated 1 month agoApache-2.0
Docker · Web#Multi-user access#Prompt versioning
Pezzo is an open-source platform for developers and teams managing prompts and monitoring LLM applications. You can run the full stack locally with Docker Compose, keeping the prompt management and monitoring platform on infrastructure you control. Its source code uses the Apache 2.0 license.
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.
3.2KUpdated 3 months agoMIT
#Guardrails
LLM Guard is a Python security toolkit for developers building applications around large language models. It checks prompts and generated responses for risks such as prompt injection, sensitive data exposure and harmful language. The project is archived and no longer maintained, including its associated models on Hugging Face.
4.6KUpdated 19 hours agoMIT
Web#AI red teaming#OpenAI-compatible API
PyRIT is an MIT-licensed, open source Python framework for security professionals and engineers assessing generative AI systems. It combines automated attack testing with human-led investigations through CoPyRIT, a web interface served locally. The framework runs locally, but prompts go to the target services you choose; cloud targets and cloud-based scorers process requests outside your machine.
1.5KUpdated 3 years agoApache-2.0
Web#Guardrails#Semantic search
Rebuff is a prompt injection detector for developers building LLM applications that accept untrusted input. It combines checks for suspicious prompts with a record of past attacks and tests for leaked prompt content. The project is archived and no longer maintained.
496Updated 3 years agoApache-2.0
Docker · Web#Guardrails#Semantic search
Vigil is a self-hosted security scanner for developers and researchers who want to check LLM inputs and responses for prompt injection, jailbreak attempts, and other suspicious content. It combines several detection methods and includes attack signatures and datasets, so teams can assess known threats without building every detector themselves. It is experimental alpha software for research and is open source under Apache 2.0.
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
2.2KUpdated 20 hours agoApache-2.0
Docker#LLM tracing#MCP#Multi-user access
Agent Router is an open source AI gateway for teams whose agents use both model APIs and MCP tools. It runs on a laptop, a dedicated gateway, or Kubernetes, and gives applications one OpenAI-compatible entry point for cloud providers and self-hosted inference. The project uses the Apache 2.0 license.
1.2KUpdated 2 years agoMIT
Docker#Guardrails#Multi-user access#OpenAI-compatible API
BricksLLM is a self-hosted AI gateway for teams that need to control access and spending across LLM applications. It sits between applications and model providers, applying cost and rate limits to individual API keys. The gateway is open source under the MIT license and runs locally or on your server through Docker, with PostgreSQL and Redis.
5.6KUpdated 2 years agoApache-2.0
#Ollama integration#OpenAI-compatible API
RouteLLM is a self-hosted Python framework for developers who want to split requests between a stronger LLM and a cheaper model. It judges which prompts need the stronger model, so an application doesn't have to send every request to its most expensive provider. You control the cost-quality tradeoff through a routing threshold.
37KUpdated 2 years agoMIT
Linux · Docker · Web#Multi-user access#Ollama integration#OpenAI-compatible API
One API is a self-hosted LLM API gateway for developers and teams that want to share model access across apps or users. It puts cloud providers and Ollama behind an OpenAI-compatible API, so clients can use one endpoint across different backends. It's open source under MIT and runs on your own server as a single executable or in Docker.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
55.1KUpdated 2 years agoMIT
Windows · Docker#Code execution#Multimodal input
gpt-engineer is a locally run coding assistant for developers who want to experiment with AI code generation and build their own agents. It can create software from plain-language descriptions or make requested changes to an existing codebase. The project is archived and no longer maintained. Its Python code is open source under the MIT license.
4.4KUpdated 20 hours ago
#Guardrails#Multilingual#Multimodal input
Llama Guard is Meta's collection of downloadable AI content moderation models for developers building LLM applications. It checks user inputs and model responses for content that violates safety policies, including text and images. Developers can use the models in their own deployments or access moderation through Meta's hosted Llama API.
4.2KUpdated 1 year agoApache-2.0
Windows · Linux · Web · VS Code#Batch processing#Guardrails#Hugging Face integration
LMQL is a programming language for developers who need model calls and ordinary Python logic in the same program. It lets you define rules for generated text, including types, length limits, allowed answers and stopping phrases. Those rules apply during generation, so you can constrain intermediate responses as well as the final output.
3.7KUpdated 1 day agoMIT
Docker · Web · VS Code#Git integration#Hybrid search#llama.cpp backend
Twinny is an AI coding assistant for VS Code that lets developers choose where their models run: on their own computer, a private server or a hosted API. It's for individuals and teams who want code suggestions and repository chat with control over where their code goes. The extension and team gateway are open source under the MIT license.
47.9KUpdated 1 day agoGPL-2.0
Linux · Web#LLM tracing#Multi-user access#Structured output
Discourse AI is the official AI plugin bundled with Discourse. Forum administrators can enable its features independently, including an AI bot, semantic search, topic and chat summaries, spam detection and writing assistance. It runs inside a Discourse community, which you can host on your own server.
5.2KUpdated 4 days agoApache-2.0
Linux · Docker · Web#Hugging Face integration#LoRA#Quantization
H2O LLM Studio is a self-hosted tool for teams that want to adapt language models to their own datasets without writing training code. Its browser interface brings training experiments, evaluation, and model testing into one place. The project is open source under Apache 2.0.
2.4KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Ollama integration
UpTrain is an open source LLM evaluation tool for developers who need to measure answer quality and investigate failures in their AI applications. Its self-hosted web dashboard runs on your machine through Docker, with a Python package for evaluations inside application code. The dashboard requires no coding.
1.1KUpdated 2 years agoApache-2.0
#OpenAI-compatible API
LLMPerf is a Python library for developers comparing LLM API performance and checking basic response correctness. The project is archived and no longer maintained. Its open-source code uses the Apache License 2.0.
953Updated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration
Ollama Grid Search is a desktop app for comparing LLM responses across models, prompts and inference settings. It runs on macOS, Windows and Linux, and suits developers or anyone choosing a model and prompt combination for a particular task. You can inspect the responses together rather than repeat each test by hand.