Experiment tracking I actually keep using
In short: I use DVC to version data, MLflow to record runs and models, and Weights and Biases when I want to watch a long training live or share charts. What makes them useful is one habit: every run…
- published
- read time
- 4 min
- words
- 782
- lang
- en
- filed under
- Engineering
In short: I use DVC to version data, MLflow to record runs and models, and Weights and Biases when I want to watch a long training live or share charts. What makes them useful is one habit: every run records the data version and the git commit it came from.
I've used all three of these tools across different projects, some alone and some on teams. Each one is good. Each one also has a way of turning a small project into an infrastructure project if you let it. This is the setup that survived, and the parts I stopped using.
Three questions a tracker has to answer
Months after a run, I only ever want to know three things about it:
- Which data did it see?
- Which code and settings produced it?
- What came out, and was it better than last time?
Most frustration with experiment tracking comes from asking one tool to answer all three. They each answer one of them well.
| Job | DVC | MLflow | Weights and Biases |
|---|---|---|---|
| Version the dataset | Main use | Possible, clumsy | Possible, via artifacts |
| Record params and metrics | Possible | Main use | Main use |
| Store and load models | Possible | Main use | Possible |
| Watch training live | No | Basic | Main use |
| Share charts with others | No | Needs a server | Main use |
| Keep everything on my machine | Yes | Yes | Not by default |
The table is my own reading of how each tool fits, not a feature list. All three do more than this. The question is what each does without a fight.
What stuck
DVC for data
DVC keeps a small text file in git that points at the real data, which lives somewhere else. The file holds a hash. Change one image in the training set and the hash changes. That's all I need from it, and it's the part I've kept everywhere.
dvc add data/train
git add data/train.dvc .gitignore
git commit -m "train set v2: relabelled masks"
dvc push
MLflow for runs
MLflow works with nothing but a local file. No server, no account. I point it at a SQLite file in the project and open the UI when I want to compare runs. When I need a model again, I load it from the run instead of hunting for a file called final_v3_really.pth.
Weights and Biases for long training
For a model that trains for hours, I want to see the loss curve from my phone and send a link to someone. That is where Weights and Biases is clearly better than the other two. I use it for those runs and not for the quick ones.
The habit that ties them together
None of the tools links data, code and results for you unless you ask. So every training script starts the same way: read the git commit, read the data hash from the .dvc file, and attach both to the run.
import subprocess
import yaml
import mlflow
commit = subprocess.check_output(["git", "rev-parse", "HEAD"], text=True).strip()
with open("data/train.dvc") as f:
data_md5 = yaml.safe_load(f)["outs"][0]["md5"]
mlflow.set_tracking_uri("sqlite:///mlflow.db")
mlflow.set_experiment("baseline")
with mlflow.start_run():
mlflow.set_tags({"git_commit": commit, "data_md5": data_md5})
mlflow.log_params({"lr": 1e-3, "batch_size": 16, "epochs": 100})
model, val_dice = train() # your training function
mlflow.log_metric("val_dice", val_dice)
With those two tags, any run can be traced back to its inputs. That is the lineage, and it's the thing I actually use.
What got in the way
- Running a tracking server for one person. A shared MLflow server makes sense for a team. For me alone, it was one more thing to keep alive. The local file is enough.
- DVC pipelines during exploration. Defining every stage in a pipeline file is great once the steps are stable. While I'm still changing the preprocessing every hour, it slows me down. I add stages at the end, not the start.
- Hosted dashboards and sensitive data. In healthcare work, the data and often the metrics shouldn't leave the building. A hosted tracker needs a careful look before it goes near that kind of project. Local MLflow doesn't.
- Logging everything. Autologging every metric at every step fills the UI with curves I never open. I log the few numbers I'd actually compare.
Monitoring a model after it ships is a different job again, with its own tools. I keep it separate from experiment tracking on purpose.
Where to start
If you track nothing today, start small. Run dvc init and dvc add on your training data. Point MLflow at a local SQLite file. Add the few lines that tag each run with the commit and the data hash. Add Weights and Biases later, when a run takes long enough that you want to watch it.
related