Writing

Explaining a model to a doctor

In short: doctors don't want to see how the model works. They want to know why it said this about this patient, in units they already use, and whether they should trust it. A simple bar chart with…

published
read time
4 min
words
886
lang
en
filed under
Engineering

In short: doctors don't want to see how the model works. They want to know why it said this about this patient, in units they already use, and whether they should trust it. A simple bar chart with plain feature names does that. Beeswarms and log-odds mostly don't.

I've used SHAP and LIME on clinical models, mostly on tabular data at a hospital network, and the same problem follows me to every new clinical project. The tools are good. The plots they produce by default are made for data scientists. Put them in front of a clinician and you can watch the room lose interest, or worse, watch someone draw the wrong conclusion with full confidence.

This is what I've seen work, and what I've stopped showing.

The question behind the question

When a doctor asks "how does the model decide?", it's rarely a request for the architecture. It's usually one of three questions:

  1. Is it looking at the right things? They want to see whether the top signals make clinical sense.
  2. Why this patient? They have a case in front of them and the score surprised them.
  3. When should I ignore it? They want to know the situations where the model is out of its depth.

Global importance answers the first. A per-patient explanation answers the second. Nothing in SHAP answers the third on its own. That one needs error analysis by subgroup, and an honest conversation.

The plot that helped

The most useful picture is also the plainest: a horizontal bar chart of mean absolute SHAP value per feature, top eight or ten, with names a clinician would use. Not an abbreviated column header but "systolic blood pressure". Not a site code but "which lab ran the test".

systolic pressure age lab that ran the test visit count medication count body mass index mean |SHAP|, higher means more influence ?
An illustrative importance chart. The third bar is the one a clinician will ask about, and they should.

This chart does two jobs. It lets a clinician agree that the top signals are plausible, which builds the right kind of trust. And it surfaces the odd one out. In the example, the lab that ran the test ranks high. No disease depends on which lab drew the blood. That bar is a question about how the data was collected, and a clinician will spot it faster than an engineer will, because they know which signs to expect.

For a single patient, the equivalent is a waterfall or a short sorted list: these three things pushed the score up, these two pulled it down, with the actual values next to them. "Systolic pressure 162, higher than most" reads like a clinical note. A force plot with arrows and a base value reads like a physics exam.

The plots that confused

Some plots caused more harm than good when I showed them to people who don't work with models every day.

What worked and what didn't

Helped

  • Bar chart of mean |SHAP|, top features only
  • Plain clinical names with units
  • Per-patient list: up, down, actual value
  • Two or three similar past cases next to the score
  • Performance split by site, device, age group

Confused

  • Beeswarm plots with a colour scale for feature value
  • Explanations in log-odds
  • Interaction plots between two features
  • LIME explanations that change when you rerun them
  • Thirty features at once

The beeswarm is the one data scientists love most. It packs direction, size and feature value into one picture. That's exactly the problem. It takes a few minutes to learn to read, and in a meeting nobody spends those minutes. People read the colour as good and bad.

Log-odds are the other trap. A SHAP value of 0.4 means nothing to someone who thinks in percentages and reference ranges. If you can, explain in probability space and say so, or skip the numbers and keep the direction and the rank.

LIME has a specific problem in this setting. It fits a small local model around one prediction using random samples, so two runs on the same patient can rank features differently. Show it once and someone will rerun it later, see a different answer, and stop trusting the whole system. If I use LIME at all now, I fix the random seed and check that the top features are stable across several runs first.

What explanation can't do

An explanation describes the model. It doesn't make the model right. A clear SHAP chart for a model trained on leaky data is a clear chart of a mistake. I've found it helps to say this out loud at the start of the meeting, so nobody walks out thinking "explainable" means "validated".

And some questions can't be answered with feature attributions at all. "Would it work on my patients?" is a question about data coverage. The honest answer is a table: here is how it does on each site, device and age group we tested, and here are the groups we haven't seen.

Before your next meeting

  • Rename every feature to what a clinician would call it, with units.
  • Make one global bar chart, top eight features, and nothing else on the slide.
  • Pick three real cases: one the model got right, one it got wrong, one borderline. Show the per-patient list for each.
  • Bring a subgroup table, and say which groups are missing.
  • Ask the clinician which bar surprised them. Then go and find out why.

related

Keep reading