
DVC connects data and model versions to the code in your Git repository, so you can reproduce a machine learning experiment with the inputs it used. It's a free, open-source tool under Apache 2.0 for individual data scientists and small projects. It runs on macOS, Windows and Linux.
Data and model files live in a cache outside Git, while Git records their version information alongside your code. You can keep artifacts on your machine or store and share them through cloud storage. Supported storage connections include Amazon S3, Azure, Google Cloud Storage, Google Drive and SSH.
Experiment tracking runs locally in your Git repository without a separate server. DVC lets you compare data, code, parameters, models and performance plots, so you can check what changed between experiments rather than rely on filenames or notes. You can also share experiments through existing Git hosting, including GitHub and GitLab, and reproduce someone else's experiment.
Its pipelines record how code and data produce model artifacts. When an input changes, DVC reruns only the affected steps, avoiding a full pipeline run for every edit. Git also versions the pipeline definitions, keeping the workflow tied to the project history.
DVC has a command-line interface and a VS Code extension. The extension brings experiment tracking and data management into the editor; the core tool handles versioning, pipelines and experiment reproduction.
Claim this page with an email at dvc.org. DVC gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find DVC?Promote it
Something wrong or outdated on this page?
5.2KUpdated 4 days agoApache-2.0
Linux · Docker · Web#Hugging Face integration#LoRA#Quantization
H2O LLM Studio is a self-hosted tool for teams that want to adapt language models to their own datasets without writing training code. Its browser interface brings training experiments, evaluation, and model testing into one place. The project is open source under Apache 2.0.
7.2KUpdated 1 month agoApache-2.0
macOS · Web#Works offline
TensorBoard is a browser-based toolkit for inspecting TensorFlow experiments on your own machine or server. It's for researchers and ML developers who need to understand training behavior, compare runs, and investigate model performance. It works entirely offline, including behind a corporate firewall or in a datacenter, so experiment data can stay within your own environment.
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
6.3KUpdated 9 months agoApache-2.0
Docker · Web
Aim is a free, open source ML experiment tracker for researchers and teams who want to keep training records on their own infrastructure. It runs in your training environment or on a self-hosted server, with Docker and Kubernetes deployment support. Its Apache 2.0 license permits use and modification.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
15.9KUpdated 1 month agoApache-2.0
Web#MLX#Multi-user access
Kubeflow is a self-hosted AI platform for teams that run machine learning workloads on Kubernetes. It brings model development, training and production workflows into a modular stack that can run on a local laptop, on-premises infrastructure or a cloud Kubernetes cluster. It's open source under Apache 2.0.
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.