Writing

Experiment tracking I actually keep using

In short: I use DVC to version data, MLflow to record runs and models, and Weights and Biases when I want to watch a long training live or share charts. What makes them useful is one habit: every run…

published
read time
4 min
words
782
lang
en
filed under
Engineering

In short: I use DVC to version data, MLflow to record runs and models, and Weights and Biases when I want to watch a long training live or share charts. What makes them useful is one habit: every run records the data version and the git commit it came from.

I've used all three of these tools across different projects, some alone and some on teams. Each one is good. Each one also has a way of turning a small project into an infrastructure project if you let it. This is the setup that survived, and the parts I stopped using.

Three questions a tracker has to answer

Months after a run, I only ever want to know three things about it:

  1. Which data did it see?
  2. Which code and settings produced it?
  3. What came out, and was it better than last time?

Most frustration with experiment tracking comes from asking one tool to answer all three. They each answer one of them well.

JobDVCMLflowWeights and Biases
Version the datasetMain usePossible, clumsyPossible, via artifacts
Record params and metricsPossibleMain useMain use
Store and load modelsPossibleMain usePossible
Watch training liveNoBasicMain use
Share charts with othersNoNeeds a serverMain use
Keep everything on my machineYesYesNot by default

The table is my own reading of how each tool fits, not a feature list. All three do more than this. The question is what each does without a fight.

What stuck

DVC for data

DVC keeps a small text file in git that points at the real data, which lives somewhere else. The file holds a hash. Change one image in the training set and the hash changes. That's all I need from it, and it's the part I've kept everywhere.

dvc add data/train
git add data/train.dvc .gitignore
git commit -m "train set v2: relabelled masks"
dvc push

MLflow for runs

MLflow works with nothing but a local file. No server, no account. I point it at a SQLite file in the project and open the UI when I want to compare runs. When I need a model again, I load it from the run instead of hunting for a file called final_v3_really.pth.

Weights and Biases for long training

For a model that trains for hours, I want to see the loss curve from my phone and send a link to someone. That is where Weights and Biases is clearly better than the other two. I use it for those runs and not for the quick ones.

The habit that ties them together

None of the tools links data, code and results for you unless you ask. So every training script starts the same way: read the git commit, read the data hash from the .dvc file, and attach both to the run.

import subprocess
import yaml
import mlflow

commit = subprocess.check_output(["git", "rev-parse", "HEAD"], text=True).strip()
with open("data/train.dvc") as f:
    data_md5 = yaml.safe_load(f)["outs"][0]["md5"]

mlflow.set_tracking_uri("sqlite:///mlflow.db")
mlflow.set_experiment("baseline")

with mlflow.start_run():
    mlflow.set_tags({"git_commit": commit, "data_md5": data_md5})
    mlflow.log_params({"lr": 1e-3, "batch_size": 16, "epochs": 100})
    model, val_dice = train()          # your training function
    mlflow.log_metric("val_dice", val_dice)

With those two tags, any run can be traced back to its inputs. That is the lineage, and it's the thing I actually use.

data md5 git commit params run model val_dice
Three inputs, one run, two outputs. If any input is missing, the run can't be reproduced.

What got in the way

  • Running a tracking server for one person. A shared MLflow server makes sense for a team. For me alone, it was one more thing to keep alive. The local file is enough.
  • DVC pipelines during exploration. Defining every stage in a pipeline file is great once the steps are stable. While I'm still changing the preprocessing every hour, it slows me down. I add stages at the end, not the start.
  • Hosted dashboards and sensitive data. In healthcare work, the data and often the metrics shouldn't leave the building. A hosted tracker needs a careful look before it goes near that kind of project. Local MLflow doesn't.
  • Logging everything. Autologging every metric at every step fills the UI with curves I never open. I log the few numbers I'd actually compare.

Monitoring a model after it ships is a different job again, with its own tools. I keep it separate from experiment tracking on purpose.

Where to start

If you track nothing today, start small. Run dvc init and dvc add on your training data. Point MLflow at a local SQLite file. Add the few lines that tag each run with the commit and the data hash. Add Weights and Biases later, when a run takes long enough that you want to watch it.

related

Keep reading