Writing

The boring checklist that saved every project

In short: before I train anything I want three things: a record of where every piece of data came from, a dumb baseline to beat, and one number the people paying for the work actually care about. None…

published
read time
4 min
words
891
lang
en
filed under
Engineering

In short: before I train anything I want three things: a record of where every piece of data came from, a dumb baseline to beat, and one number the people paying for the work actually care about. None of it is clever. It is the part that keeps a project alive when the model is no longer the interesting bit.

It is the last day of the year, which is a good day to write down the things I keep relearning. Over the last few years I have worked on clinical data shared across a research consortium, on consumer hardware prototypes at a large company, and on a speech-driven companion for patients. Very different work. The projects that went well had the same three things in place early, and the ones that wobbled were missing at least one of them.

  1. Day one: lineageWrite down where the data came from, which version you have, and every step between the source and your training set.
  2. Week one: a baselineScore the simplest thing that could work, on the same split you will use for the real model.
  3. Before the first demo: one metricAgree with the people who own the problem on the single number that means "better".
  4. Before shipping: check all three againThe data has changed, the baseline may have moved, and the metric may no longer be the one anyone looks at.

1. Lineage: can you walk back from any result?

Sooner or later someone points at a result and asks "where did this come from?" With clinical data the question is not optional. If a site withdraws consent, or a column turns out to be mislabelled, you need to know which models saw it.

Lineage sounds like a tool you buy. Mostly it is a habit. Every step from the raw export to the result gets a name, an input, an output, and a version you can point to.

raw export cleaned features model result 2024-03 dump clean.py v2 hash 9f1c run 14 can you walk back?
If any box in the chain has no version, the arrow at the bottom breaks there.

In practice that means the raw data is never edited in place, each transform is a script rather than a notebook cell someone ran once, and the model run records the exact data version it trained on. Tools like DVC and MLflow make this cheaper, but a folder naming rule and a text file of hashes is already most of the value.

2. A baseline before the model

A model score on its own means nothing. 0.85 is great if the dumb answer gets 0.50, and embarrassing if the dumb answer gets 0.88. On imbalanced clinical data the dumb answer is often surprisingly good.

So before any real model, I score the simplest thing that could work:

  • Predict the most common class every time.
  • Predict tomorrow's value as today's value.
  • A logistic regression on the two or three features a domain expert would name first.

The baseline does two jobs. It tells you how much room there is to improve. And it gives you a working pipeline end to end, from data to score, in the first week, which shakes out most of the boring bugs before they can hide inside a deep model.

3. One number the business cares about

Data scientists like F1. The people who own the problem rarely do. They care about something they can feel: how long a device takes to respond, whether patients keep talking to the companion, how many cases a clinician can review in an afternoon.

The numbers I am still asked about years later are of that kind. On the hardware prototypes, it was a cut in latency of about a fifth on one project. On the patient companion, it was a rise of more than half in how much people interacted with it after we improved speech recognition and text-to-speech. Nobody has ever asked me for the validation loss.

Pick one such number with the people who own the problem, write it down, and check that your model metric moves when it moves. If the two disagree, believe the business number and go find out why.

TipPut the three answers at the top of the project README, above the install steps: where the data comes from, what the baseline scores, and which number means "better". New people read the top of the README. They do not read the wiki.

Why it gets skipped

All three feel like overhead on day one, when the exciting part is the model. And all three are cheap on day one and expensive later. Reconstructing lineage after six months means asking people what they did, and they do not remember. A baseline added after the model looks like you are moving the goalposts. A metric agreed after the demo is a negotiation instead of a decision.

None of this makes a project succeed on its own. It makes failure visible early, while there is still time to change direction.

Your first hour on the next project

Open a blank file and answer three questions in a sentence each. Where exactly does the data come from, and how would I get this same version again? What does the dumbest reasonable answer score? Which one number will the owner of this problem look at? If you cannot answer one of them, that is your first task, before any model.

related

Keep reading