Favicon of MLflow

MLflow

Self-hosted AI engineering platform for tracing agents, evaluating LLMs, and tracking models. Apache 2.0 licensed, with Ollama support.

Screenshot of MLflow website

MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.

Tracing uses OpenTelemetry to record what happens inside an application or agent. You can inspect behavior and monitor quality, latency, cost, and safety. Evaluations combine built-in metrics, LLM judges, and custom checks to track quality over time and catch regressions before deployment. AI analysis can also flag problems in recorded traces.

Prompt management keeps versions and their history together, with tools to test, deploy, and automatically optimize prompts. The AI Gateway provides an OpenAI-compatible interface across providers, with request routing, rate limits, fallbacks, and cost controls. Agent Server hosts agents with streaming responses, request validation, and built-in tracing.

For model training, MLflow records parameters, metrics, and evaluation results across experiments. Its model registry supports collaborative model management, and deployment tools cover batch and live prediction workloads.

Integrations include Ollama, LangChain, LangGraph, LlamaIndex, and PyTorch, alongside cloud providers such as OpenAI and Anthropic. MLflow itself can be self-hosted, including through Docker; calls to cloud model providers still go to those services. Managed hosting is also available through Databricks, Amazon SageMaker, and Azure ML.

Similar to MLflow