Player not loading? Watch on YouTube
This MLOps lesson explains how DVC works with Git to version machine learning datasets. The instructor connects data versioning to reproducibility: changes in feature engineering can alter the training data as well as the code, so a project needs a record of both.
The demonstration uses Windows, VS Code and Git Bash. It covers initializing DVC, adding a dataset folder and committing the resulting metadata to Git. An early attempt encounters an error because Git already tracks the data. The instructor then demonstrates a separate dataset folder, where Git tracks the .dvc file and ignore rules rather than the dataset contents.
After editing a data file, the instructor checks dvc status and runs dvc add again to update its metadata. The examples also cover adding a CSV file and deleting rows. To restore an earlier dataset, the demonstration uses git checkout followed by dvc checkout. Switching the Git commit alone leaves the working data unchanged in the second example.
The lesson examines DVC's local cache and MD5 hashes to explain where the demonstrated data versions reside. Cloud storage with AWS or GCP is planned for a later session; this tutorial keeps the dataset cache on the local machine. The remaining discussion outlines future MLOps lessons rather than demonstrating automated training or deployment.