8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
756Updated 5 hours agoMIT
macOS · Windows · Linux#LLM tracing#MCP
agentacct is a local dashboard for developers who want to review what their coding agents did across projects and compare usage with recorded results. It brings task activity, file changes and test outcomes together with token counts and estimated costs. Claude Code and Codex have the fullest task records; it also supports OpenCode, Kimi Code and Hermes, with coverage varying by capability.
29.8KUpdated 1 day agoApache-2.0
macOS · Docker · Web#Code execution#LLM tracing#Multi-user access
Sim is an AI agent workspace for teams that need to build automations and control how people use them across an organization. You can self-host it on your own server or cloud with Docker or Kubernetes. Its open-source code uses the Apache 2.0 license.
28.4KUpdated 21 hours ago
Web#Human approval#LLM tracing#MCP
Mastra is a TypeScript framework for developers building AI agents and applications on their own servers or inside existing web apps. Its server runs locally or as a standalone deployment; Mastra Cloud provides a hosted alternative. Model routing connects to providers such as OpenAI, Anthropic, and Gemini, so running the framework locally doesn't keep those model requests on your machine.
157.6KUpdated 3 hours ago
Docker · Web#Code execution#MCP#OpenAI-compatible API
Dify is a source-available platform for teams building AI agents and apps on a visual canvas. Its Community Edition runs on your own server with Docker. Dify also offers a hosted cloud service, while Enterprise deployments can run in a VPC or on a self-hosted server. The Community Edition uses a custom Apache 2.0 derivative license.
22.3KUpdated 1 day agoApache-2.0
Web#Guardrails#LLM tracing#Prompt versioning
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
35.2KUpdated 2 hours ago
Docker#Agent Skills#LLM tracing#MCP
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.
59.9KUpdated 57 minutes ago
#Guardrails#MCP#Multi-user access
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.
3.3KUpdated 1 month agoApache-2.0
Docker · Web#Multi-user access#Prompt versioning
Pezzo is an open-source platform for developers and teams managing prompts and monitoring LLM applications. You can run the full stack locally with Docker Compose, keeping the prompt management and monitoring platform on infrastructure you control. Its source code uses the Apache 2.0 license.
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
2.2KUpdated 20 hours agoApache-2.0
Docker#LLM tracing#MCP#Multi-user access
Agent Router is an open source AI gateway for teams whose agents use both model APIs and MCP tools. It runs on a laptop, a dedicated gateway, or Kubernetes, and gives applications one OpenAI-compatible entry point for cloud providers and self-hosted inference. The project uses the Apache 2.0 license.
1.2KUpdated 2 years agoMIT
Docker#Guardrails#Multi-user access#OpenAI-compatible API
BricksLLM is a self-hosted AI gateway for teams that need to control access and spending across LLM applications. It sits between applications and model providers, applying cost and rate limits to individual API keys. The gateway is open source under the MIT license and runs locally or on your server through Docker, with PostgreSQL and Redis.
1.2KUpdated 10 months agoAGPL-3.0
Docker · Web#LLM tracing#Ollama integration
Langtrace is an open-source observability tool for developers who need to debug LLM applications and track their performance. You can run it locally or on your own servers with Docker and Docker Compose. Its traces follow OpenTelemetry standards.
lunary.aiLLM Evaluation and Testing
Docker · Web#LLM tracing#Multi-user access#Prompt versioning
Lunary is a self-hosted LLM observability and prompt management platform for teams building chatbots and AI agents. It brings production traces, user conversations and prompt versions into one place so developers can investigate errors and teams can assess response quality. You can host it in your own infrastructure or use its cloud service.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
7.1KUpdated 2 days agoApache-2.0
Docker#Guardrails#LLM tracing#Multi-agent workflows
Plano, formerly Arch Gateway, is a self-hosted AI gateway for developers building applications with multiple agents or model providers. It puts routing, guardrails and request tracing in a separate service, so each agent doesn't need its own implementation of that infrastructure. It's open source under Apache 2.0.
11.2KUpdated 5 months agoMIT
VS Code#Code execution#LLM tracing#Visual workflows
Prompt flow is an MIT-licensed, open-source toolkit for developers who build LLM applications and need to test their behavior before deployment. Its development tools run locally, while an optional cloud version in Azure AI supports team collaboration. Feature development has ended.
3.3KUpdated 2 weeks agoApache-2.0
Docker · Web#LLM tracing#MCP
Laminar is an open-source platform for developers who need to see why an AI agent failed and check whether a fix worked. You can self-host it with Docker or on Kubernetes, including AWS and GCP, or use its managed cloud service. It uses the Apache 2.0 license.
44.2KUpdated 2 days agoApache-2.0
#MCP
Kong Gateway puts API, LLM and MCP traffic behind a shared gateway on your own infrastructure. It's for platform teams that need consistent access controls and traffic policies across services and AI applications. The open-source gateway uses the Apache 2.0 license and runs natively on Kubernetes through Kong's official Ingress Controller.
2.6KUpdated 20 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
7.5KUpdated 1 day agoApache-2.0
#LLM tracing#Ollama integration
OpenLLMetry adds LLM tracing to the OpenTelemetry monitoring stack a team already uses. It's for developers who need to follow model calls alongside database activity and API requests in their AI applications. The extensions run within your application and send standard OpenTelemetry data to your chosen monitoring destination.
37.5KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input
Claude Code Router is an open-source local model gateway for developers who use coding agents and want to manage their model providers in one place. It runs on macOS, Windows and Linux, with Docker and a CLI with a browser interface also available. The project uses the MIT license.
21.3KUpdated 9 months agoApache-2.0
Web#Guardrails#LLM tracing#MCP
Rasa is an AI agent platform for product teams building customer-facing text and voice assistants. Teams can deploy agents on their own infrastructure and choose their models and data arrangements. Its CALM engine combines language model understanding with business flows whose code enforces rules, so an assistant can handle conversational wording while following defined processes.
3.6KUpdated 1 day agoMIT
#Batch processing#LLM tracing#MCP
TruLens is an open-source Python tool for developers who need to find why an AI agent gives a wrong answer or spends too much on a task. It pairs step-level traces with evaluation scores, so you can connect failures to retrieval, reasoning or tool calls. It uses the MIT license and can write results to a database you run.
4.7KUpdated 2 days agoMIT
Docker · Web#Git integration#LLM tracing#MCP
Latitude is a self-hosted AI agent monitoring platform for teams that need to find production failures and check that fixes work. It connects recurring problems to the sessions that caused them, then can dispatch Claude Code or Cursor with enough context to make a fix and open a pull request.
7.9KUpdated 3 weeks agoApache-2.0
Web
Evidently is an Apache 2.0 Python library for evaluating, testing and monitoring ML and LLM systems. It works with tabular and text data, including predictive models and RAG applications. You can run one-off evaluations or self-host its open-source monitoring UI. Evidently Cloud is a separate hosted service.
6.2KUpdated 2 weeks agoApache-2.0
Web#LLM tracing#Ollama integration#Prompt versioning
Helicone combines an AI gateway with LLM observability for engineers building agents, chatbots, and document processing apps. You can self-host the open source observability platform under Apache 2.0 or use the hosted service. It also integrates with Ollama for apps that run models locally. Its hosted gateway routes requests to external AI providers, so those requests leave your machine.
18.5KUpdated 23 hours agoApache-2.0
#LLM tracing#Multimodal input
DeepEval is a Python framework for testing AI agents, RAG pipelines, and chatbots in your own environment. It's for developers and ML teams who need to compare models or prompts and catch quality regressions before deployment. The open-source framework uses the Apache 2.0 license and fits into Pytest, Python scripts, notebooks, and CI/CD.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.