Pezzo is an open-source platform for developers and teams managing prompts and monitoring LLM applications. You can run the full stack locally with Docker Compose, keeping the prompt management and monitoring platform on infrastructure you control. Its source code uses the Apache 2.0 license.
Prompt design, version management and collaboration share one console. Teams can manage prompts centrally and deliver prompt changes to their applications, with monitoring and troubleshooting tools alongside them. This makes it relevant to projects where prompts need ongoing maintenance and several people contribute to AI behavior.
Pezzo supports Node.js, Python and LangChain clients. Each supports prompt management, observability and caching, so the choice of client doesn't restrict those core capabilities. Caching can reduce costs and latency, while monitoring helps teams investigate issues in their AI operations.
The platform has a server and a browser-accessible console. Its infrastructure uses PostgreSQL, ClickHouse, Redis and Supertokens, all open-source technologies that can run through Docker Compose. Pezzo also supports OpenAI's GPT-4o. Requests to OpenAI use that external model service even when Pezzo itself is self-hosted.
Claim this page and we'll verify you by hand. Pezzo gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Pezzo?Promote it
Something wrong or outdated on this page?
lunary.aiLLM Evaluation and Testing
Docker · Web#LLM tracing#Multi-user access#Prompt versioning
Lunary is a self-hosted LLM observability and prompt management platform for teams building chatbots and AI agents. It brings production traces, user conversations and prompt versions into one place so developers can investigate errors and teams can assess response quality. You can host it in your own infrastructure or use its cloud service.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
6.2KUpdated 2 weeks agoApache-2.0
Web#LLM tracing#Ollama integration#Prompt versioning
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
Helicone combines an AI gateway with LLM observability for engineers building agents, chatbots, and document processing apps. You can self-host the open source observability platform under Apache 2.0 or use the hosted service. It also integrates with Ollama for apps that run models locally. Its hosted gateway routes requests to external AI providers, so those requests leave your machine.