
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
Tracing uses OpenTelemetry to record what happens inside an application or agent. You can inspect behavior and monitor quality, latency, cost, and safety. Evaluations combine built-in metrics, LLM judges, and custom checks to track quality over time and catch regressions before deployment. AI analysis can also flag problems in recorded traces.
Prompt management keeps versions and their history together, with tools to test, deploy, and automatically optimize prompts. The AI Gateway provides an OpenAI-compatible interface across providers, with request routing, rate limits, fallbacks, and cost controls. Agent Server hosts agents with streaming responses, request validation, and built-in tracing.
For model training, MLflow records parameters, metrics, and evaluation results across experiments. Its model registry supports collaborative model management, and deployment tools cover batch and live prediction workloads.
Integrations include Ollama, LangChain, LangGraph, LlamaIndex, and PyTorch, alongside cloud providers such as OpenAI and Anthropic. MLflow itself can be self-hosted, including through Docker; calls to cloud model providers still go to those services. Managed hosting is also available through Databricks, Amazon SageMaker, and Azure ML.
Claim this page with an email at mlflow.org. MLflow gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MLflow?Promote it
Something wrong or outdated on this page?
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
lunary.aiLLM Evaluation and Testing
Docker · Web#LLM tracing#Multi-user access#Prompt versioning
35.2KUpdated 2 hours ago
Docker#Agent Skills#LLM tracing#MCP
22.3KUpdated 1 day agoApache-2.0
Web#Guardrails#LLM tracing#Prompt versioning
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
Lunary is a self-hosted LLM observability and prompt management platform for teams building chatbots and AI agents. It brings production traces, user conversations and prompt versions into one place so developers can investigate errors and teams can assess response quality. You can host it in your own infrastructure or use its cloud service.
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.