Writing

Most AI never leaves the notebook

In short: most models die between the demo and the first real user, and rarely because the model was bad. They die because nobody owns them, the real data looks different, the metric doesn't match a…

published
read time
5 min
words
968
lang
en
filed under
Product

In short: most models die between the demo and the first real user, and rarely because the model was bad. They die because nobody owns them, the real data looks different, the metric doesn't match a decision, the data can't move, or the thing is too slow where it has to run.

I've built models in a hospital network, in a research consortium, in a hardware company's research lab and, more recently, on clinical audio. In every one of those places I've seen more notebooks than products. Some of that is healthy. Research is supposed to try things that don't work out. But a lot of good work dies for reasons that have nothing to do with the model, and those reasons repeat.

notebook demo pilot in use
The shape, not a measurement. Each step loses most of what entered it.

Five reasons, from what I have seen

1. Nobody owns it after the demo

At the hardware company, a big part of my role was prototyping new features for consumer devices, demoing them, and judging whether they were worth integrating into a product. That is a filter by design. Most prototypes shouldn't ship. But I also saw good ones stall, and the reason was usually the same: the demo went well, everyone nodded, and then there was no team whose job it was to take it further.

A model needs an owner the day after the demo: someone who will be paged when it breaks and who has time in their plan to maintain it. If you can't name that person, the model is a demo, however good it is.

2. The real data isn't the notebook data

The notebook dataset was cleaned by hand, once, by the person who built the model. Production data arrives raw, from devices and people you haven't met. In audio this is brutal: a new microphone, a new room, a different app version, and the inputs shift under the model without anyone noticing.

The fix isn't a better model. It's moving the cleaning out of notebook cells into a function that runs the same way in training and in production, plus monitoring that tells you when the inputs drift.

3. The metric doesn't match a decision

An accuracy number is not a use. In clinical work the questions are: who looks at this score, at what moment, and what do they do differently because of it? If the score arrives after the decision has been made, or flags so many patients that nobody can follow up, it doesn't matter how good the curve is. I've learned to ask clinicians these questions before training, not after.

4. The data can't go where the model is

In healthcare the data often cannot leave the building, and for good reason. In the consortium I worked with, several institutions each held their own data, under their own rules. A model that only works on one central copy of the data may never get the chance to exist in a setup like that. That is why part of my work there was federated learning, and why securing the models themselves mattered as much as training them.

Privacy and security review is not an obstacle to plan around at the end. It's a design input from day one.

5. It's too slow or too costly where it has to run

A model that answers in a few seconds on a workstation GPU can be useless on a small device, or on a scaled-to-zero cloud service that has to load it first. On consumer device features, an important part of my work was latency, not accuracy. On a virtual companion for patients, the gains came from fixing speech recognition and text-to-speech, the parts around the language model, not the language model itself.

0
raw patient rows moved between sites in the federated setup
about a fifth
lower latency, on one consumer device project
over half
more user interaction on one patient companion, after fixing speech

None of those three numbers is about model accuracy. All of them are about whether a model could actually be used.

A path out of the notebook

When I built a respiratory sound classifier, I had a short amount of time and tried to build it the way I'd build something meant to leave the notebook. The training is a script with arguments, not cells. There is a debug mode that runs the whole pipeline on random data. Every run is logged in MLflow. GitHub Actions runs a small CI pipeline. A Streamlit app serves it, Prometheus and Grafana collect metrics, and there is a public demo on Hugging Face. None of that is advanced. It's just the order of steps that keeps showing up.

  1. ScriptMove the notebook into a script with arguments, and a debug mode that runs end to end on fake data.
  2. TrackLog every run, its data version and its settings, so you can say which model is which.
  3. PackagePut the model and its preprocessing in one container. Same code path for training and serving.
  4. ServeAn endpoint or a small app that a real user can try, behind CI that runs on every push.
  5. WatchMetrics for latency, errors and input drift, and a person who looks at them every week.
  6. Hand overA named owner, a short runbook, and a date for the first review.

Pick one model this week

Take one model you've built that is still in a notebook. Write down, in one line each, the answers to five questions: who owns it, what data it will see, which decision it changes, where the data is allowed to be, and how fast it must answer on the target hardware. Any line you can't fill in is the reason it hasn't shipped. Start there, not with a new architecture.

related

Keep reading