The clinician said no, and was right
In short: a clinical model can score beautifully because it has learned something the clinic already decided. The fastest way to catch that is not a better metric. It is ten minutes with the person…
- published
- read time
- 5 min
- words
- 955
- lang
- en
- filed under
- Engineering
In short: a clinical model can score beautifully because it has learned something the clinic already decided. The fastest way to catch that is not a better metric. It is ten minutes with the person who fills in the data.
I've just moved on from years of work inside a hospital network. Before new work fills my head, I want to write down the lesson from the hospital years that I think about most. It isn't about architecture. It's about a meeting where I showed a number I was proud of, and a clinician looked at it and said no.
I'll keep the details blurred, because the project and the people deserve that. The shape of it is what matters, and the shape is common.
The number that looked great
In one project we had a tabular model on clinical records. The task was to flag patients who might need a closer look. The features were the usual mix: demographics, scores from questionnaires, visit history, a few coded fields from the record system.
The model did very well. Not suspiciously well at first glance, just clearly better than the simple baseline. The scores for the two groups barely overlapped. Cross-validation agreed with itself. I made the plots, wrote the summary, and brought it to the team.
The clinician on the team looked at the feature importance list before looking at the metric. Near the top was a coded field that I had treated as just another column. The question was simple: when does that field get filled in? The answer came a second later, from the same person: after the team has already decided the patient needs follow-up. It was not a cause and it was not an early sign. It was a record of the decision we were trying to predict.
The model wasn't predicting the outcome. It was reading the clinic's own notes about it.
What the plots looked like before and after
Once we dropped that field, and every other field that is only known after the decision, the picture changed. The two groups still separated, but much less. The model went from impressive to modest. Modest was the honest version.
This is called target leakage, and every ML course mentions it. What the course doesn't tell you is that in clinical data you usually can't see it from the data alone. The field had a neutral name. Its values were plausible. Nothing in the distribution said "I was written after the fact". Only someone who works the process every day knows the order in which things happen.
Why the engineer can't catch this alone
A record system stores the state of the world at export time, not at the moment a decision was made. That flattens time. A diagnosis code, a referral, a medication, a booked appointment, a note template: all of them sit in the same table as the questionnaire answers, with no sign of which came first.
The usual ways this shows up:
- Fields written after the decision. Referrals, follow-up flags, treatment codes. They predict the label because they are the label, delayed.
- Process signals. Patients who are worried about get more visits, more tests and longer notes. Count of visits becomes a proxy for "the team was already concerned".
- Site and workflow. One clinic fills in a form that another skips. The model learns which clinic you came from.
- Repeated patients. The same person shows up in training and test as different rows, and the model recognises them instead of the condition.
The engineer can check for the last one with a grouped split. The first three need someone who knows the workflow.
What I do differently now
I stopped bringing the metric first. I bring the feature list first, and for each of the top features I ask one question: at the moment you would use this model, do you already know this value? If the answer is "no" or "it depends", the feature goes out until we understand it.
I also ask clinicians to look at a handful of individual cases the model got very right. Strangely, those are more useful than the errors. A case the model nails with high confidence, where the clinician says "that one was obvious, we'd already referred her", is a red flag. The model is being rewarded for agreeing with a decision that was already made.
The part that's hard to say out loud
The modest model was worth less in a slide and worth more in a clinic. A score that tells you what you already decided has no value at the bedside. A weaker score that is available earlier, before anyone has decided, might.
That meeting was uncomfortable. I had been happy with a number, and it was taken apart in two sentences. But the clinician was right, and saved us from building something that would have looked fine in a paper and done nothing for a patient.
Try this on your own project
- Sort your features by importance and take the top ten.
- Sit with someone who works the process and ask, for each one, when it gets written and by whom.
- Drop anything written at or after the decision, then retrain with a split grouped by patient.
- Show them five cases the model was most confident about, and listen for "we already knew".
If the metric drops, good. You've just found out what your model actually knows.
related