LLMOps: Gateways, Observability and Guardrails

Self-hosted tools for LLM apps in production, including gateways that route between models, tracing, evals and guardrails.

Subcategories

98 tools
Open-source LLM observability software with self-hosting via Docker, OpenTelemetry traces, and Python and TypeScript SDKs.

1.2KUpdated 10 months agoAGPL-3.0

Docker · Web#LLM tracing#Ollama integration

Self-hosted LLM observability platform for agent tracing, chatbot analytics and prompt management, with Docker, Kubernetes and a cloud offering.

lunary.aiLLM Evaluation and Testing

Docker · Web#LLM tracing#Multi-user access#Prompt versioning

Self-hosted LLM gateway with evaluation, A/B testing, and Ollama support. Open source under Apache 2.0; archived and no longer maintained.

11.7KUpdated 4 months agoApache-2.0

Docker · Web#Batch processing#LLM tracing#Multimodal input

Favicon of AutoGen

AutoGen

2 videos
AI agent framework for Python with local and distributed runtimes, a browser-based prototyping UI, and OpenAI and Azure OpenAI integrations.

61.2KUpdated 6 months agoCC-BY-4.0

Web#Code execution#MCP#Multi-agent workflows

Self-hosted AI application server with an OpenAI-compatible API, local Ollama and vLLM backends, document search and agent tool calling. MIT licensed.

8.4KUpdated 21 hours agoMIT

#Agent Skills#Batch processing#Guardrails

Self-hosted AI agent framework in C# for Windows, Linux and macOS, with LLamaSharp and cloud provider plugins. Apache 2.0 licensed.

3.1KUpdated 2 days agoApache-2.0

macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows

A self-hosted AI serving framework that deploys models and pipelines on Kubernetes, on-premises or in the cloud, under the Business Source License.

4.8KUpdated 8 months ago

An open source LLM evaluation framework for local models and hosted APIs, with GGUF, Hugging Face transformers and llama.cpp support. MIT licensed.

14.1KUpdated 2 weeks agoMIT

macOS#Batch processing#GGUF#Hugging Face integration

An open-source AI gateway built on Envoy. Self-host agent orchestration, route LLM requests, and capture OpenTelemetry traces under Apache 2.0.

7.1KUpdated 2 days agoApache-2.0

Docker#Guardrails#LLM tracing#Multi-agent workflows

An open-source AI gateway under the MIT license, with an OpenAI-compatible API for Ollama and cloud providers. Runs locally with Node.js or Docker.

13.1KUpdated 4 months agoMIT

Docker · Web#Guardrails#Ollama integration#OpenAI-compatible API

An open-source toolkit for building and evaluating LLM apps locally, with Python workflows, a VS Code designer and optional Azure AI collaboration.

11.2KUpdated 5 months agoMIT

VS Code#Code execution#LLM tracing#Visual workflows

An open-source LLM evaluation platform under Apache 2.0 that compares Hugging Face models and cloud APIs across reasoning, coding, safety and other tasks.

7.5KUpdated 2 days agoApache-2.0

#Distributed execution#Hugging Face integration#OpenAI-compatible API

An open-source AI agent observability platform with Docker self-hosting, Apache 2.0 licensing, automated failure detection, and regression evals.

3.3KUpdated 2 weeks agoApache-2.0

Docker · Web#LLM tracing#MCP

Self-hosted API and AI gateway with Apache 2.0 licensing, Kubernetes support, and routing across OpenAI, Anthropic, Gemini and other LLM providers.

44.2KUpdated 2 days agoApache-2.0

#MCP

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

2.6KUpdated 21 hours agoApache-2.0

Web#OpenAI-compatible API#Prompt caching

An open-source Python library for AI agent evaluation and red teaming, licensed under Apache 2.0, with a hosted or on-premise security platform.

5.8KUpdated 1 day agoApache-2.0

#AI red teaming

An open-source Python toolkit for evaluating embedding models and retrieval systems, with Sentence Transformers support and an interactive leaderboard.

3.4KUpdated 22 hours agoApache-2.0

Docker · Web#Multilingual#Multimodal input

Self-hosted LLM observability and evaluation platform under Apache 2.0. Monitor Ollama, vLLM and coding agents with portable OpenTelemetry traces.

2.8KUpdated 1 day agoApache-2.0

Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning

A self-hosted AI gateway under Apache 2.0 that runs in Docker, manages LLM APIs and MCP servers, and supports Kubernetes inference routing.

9.5KUpdated 1 day agoApache-2.0

Docker · Web#Guardrails#MCP#Tool calling

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 21 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

An open-source LLM training and deployment platform under Apache 2.0. Build specialized models on your own infrastructure or use its hosted service.

9.4KUpdated 2 days agoApache-2.0

Docker#Distributed execution#LoRA#Multimodal input

An open-source local LLM stack manager that connects Ollama, llama.cpp and AI apps through Docker Compose. Includes a CLI and companion app.

3.2KUpdated 5 days agoApache-2.0

macOS · Linux · Docker#GGUF#llama.cpp backend#MCP

AI model hub with an Apache 2.0 Python library for local inference, training and evaluation, plus hosted demos and cloud notebooks.

9.2KUpdated 7 days agoApache-2.0

Docker#Image-to-image#Inpainting#Multimodal input

Open-source LLM observability for Python and JavaScript/TypeScript apps, with Ollama support and standard OpenTelemetry output. Licensed under Apache 2.0.

7.5KUpdated 1 day agoApache-2.0

#LLM tracing#Ollama integration

A local model gateway that routes coding agents to your chosen providers. Runs on macOS, Windows, Linux or Docker and uses the MIT license.

37.5KUpdated 4 days agoMIT

macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input

Open-source LLM training toolkit built on PyTorch. Train and chat with your own models on NVIDIA GPUs, with smaller CPU and Apple Silicon examples.

58.3KUpdated 3 months agoMIT

macOS#Code execution#Tool calling

An open-source MCP gateway that runs tool servers in Docker and connects Claude Code, Cursor and Zed through shared profiles. Licensed under MIT.

1.6KUpdated 2 weeks agoMIT

#MCP#Tool calling

An open-source Python framework for testing LLMs and AI agents locally, with adversarial attacks, production guardrails, and Apache 2.0 licensing.

3KUpdated 1 week agoApache-2.0

#AI red teaming#Guardrails

A self-hosted AI agent platform for text and voice, with business rules, conversation testing, and a shared visual and code workspace.

21.3KUpdated 9 months agoApache-2.0

Web#Guardrails#LLM tracing#MCP

Favicon of ECC

ECC

2 videos
A coding assistant toolkit for Claude Code and Codex, with local Markdown memory, MIT-licensed workflows, and support for self-hosted model endpoints.

269.7KUpdated 21 hours agoMIT

Windows#Guardrails#MCP#Multi-agent workflows

An open-source Python tool for AI agent evaluation and tracing, with MIT licensing, OpenTelemetry support and a database you control.

3.6KUpdated 1 day agoMIT

#Batch processing#LLM tracing#MCP

An open-source AI agent monitoring platform under MIT. Self-host with Docker or Kubernetes, track failures, and dispatch Claude Code or Cursor to fix them.

4.7KUpdated 2 days agoMIT

Docker · Web#Git integration#LLM tracing#MCP

A Python framework for generating and evaluating LLM datasets, with Apache 2.0 licensing and integrations for Anthropic, Cohere and Argilla.

3.4KUpdated 10 months agoApache-2.0

#Structured output

Self-hosted AI gateway connects apps to OpenAI, Anthropic and Gemini APIs, with provider routing, access controls and usage tracking. Open source under AGPL-3.0

49.1KUpdated 2 days agoAGPL-3.0

Linux · Docker · Web#Multi-user access#Multimodal input#OpenAI-compatible API

Open-source Python framework for LLM evaluation, RAG testing and ML monitoring, with a self-hosted dashboard and an Apache 2.0 license.

7.9KUpdated 3 weeks agoApache-2.0

Web

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 24 hours agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API