Writing

Monitoring a model that listens to patients

In short: for a voice model in healthcare, the true labels arrive late or never, so you can't watch accuracy. You watch the recordings coming in, the scores going out, and the plumbing in between, and…

published
read time
4 min
words
898
lang
en
filed under
Engineering

In short: for a voice model in healthcare, the true labels arrive late or never, so you can't watch accuracy. You watch the recordings coming in, the scores going out, and the plumbing in between, and you compare each week against a reference period you trust.

A model that scores voice recordings fails quietly. It doesn't crash when a new phone model shows up, or when a clinic moves its recording station next to a noisy corridor. It keeps returning numbers. The numbers just stop meaning what they meant at validation time. Monitoring is how you notice before a clinician does.

Why you can't just track accuracy

In most product ML you get feedback quickly: a click, a purchase, a correction. In clinical work the confirmation of a diagnosis can take weeks or months, comes through a different system, and often never makes it back to you at all. By the time you could compute accuracy on last month's traffic, last month's patients have already been scored.

So the useful signals are the ones you have at prediction time:

  • What comes in. Duration, sample rate, loudness, how much of the clip is silence, how much is clipped, which device and app version sent it.
  • What goes out. The distribution of scores, and the share of recordings above the decision threshold.
  • The plumbing. Latency, error rate, how many uploads failed or were rejected as unusable.

None of these tells you the model is right. Together they tell you whether the world still looks like the one the model was validated on.

Drift, in one picture

I use Evidently for the drift part. You give it a reference dataset (a period where you trust the inputs, often the validation set) and a current window, and it runs a statistical test per feature and reports which ones have moved. The number I chart week to week is the share of features that drifted.

0 alert weeks new app version alert fires
An illustrative week-by-week drift line. The jump lines up with a change on the recording side, not in the patients.

The picture above is the common case. Drift in medical audio is far more often about how the audio was captured than about who was captured. An app update changes the default gain. A new tablet has a different microphone. A site starts recording at a different sample rate. The patients didn't change. The pipeline did.

The alerts I'd actually keep

Too many alerts and people mute the channel. These are the ones worth a message, with the first thing to check when each fires. The conditions are starting points, tune them on your own reference period.

AlertFires whenFirst thing to check
Input driftShare of drifted features crosses the line for two windows in a rowDevice and app version mix for the same window
Unusable recordingsShare rejected for silence or clipping climbsOne site or one device behind most of the rejects
Score shiftShare above the decision threshold moves well outside its usual rangeWhether input drift fired too; if not, the model or its config changed
Missing dataA site that usually sends recordings sends noneUpload errors, then call the site
Latency and errorsSlow or failed requests riseCold starts, instance limits, a bad deploy

The score shift alert is the one that makes people nervous, and it should. If the share of patients flagged doubles, someone downstream is getting twice the referrals. Even if the model is technically right, that is a conversation to have before it happens, not after.

The weekly job

The drift check runs on a schedule, once a week, over the features extracted from that week's recordings. It writes an HTML report you can open and a small summary that the alerting reads. This uses the Evidently report API as it is in the version I'm on; the library moves fast, so check the docs for yours.

import pandas as pd
from evidently.report import Report
from evidently.metrics import DatasetDriftMetric

FEATURES = ["duration_s", "rms_db", "silence_ratio", "clip_ratio",
            "pitch_mean", "pitch_std", "score"]

ref = pd.read_parquet("reference/features.parquet")[FEATURES]
cur = pd.read_parquet("weekly/2025-W14.parquet")[FEATURES]

report = Report(metrics=[DatasetDriftMetric()])
report.run(reference_data=ref, current_data=cur)
report.save_html("reports/drift_2025-W14.html")

result = report.as_dict()["metrics"][0]["result"]
summary = {
    "week": "2025-W14",
    "drift_share": result["share_of_drifted_columns"],
    "dataset_drift": result["dataset_drift"],
    "n_current": len(cur),
}
print(summary)  # the alerting job reads this

Two details matter more than the code. The features have to be computed by the same function in training and in production, or you are measuring your own bug. And the reference set has to be one you trust, which means you pick it on purpose and keep it fixed until you deliberately retrain.

What I look at every week

Alerts catch the loud problems. A short weekly look catches the slow ones. Mine is a fifteen-minute checklist:

  1. Open the drift report and read the top three moved features, even if no alert fired.
  2. Check the score histogram against last week and against the reference.
  3. Look at volume per site. A quiet site is often a broken site.
  4. Listen to a handful of recordings that were rejected, and a handful that scored highest.
  5. Write one line in a log: what changed, and whether anyone needs to know.

Step four sounds old-fashioned. It's the most useful one. Ten seconds of a recording tells you about a new background noise faster than any statistical test. If you run a model on audio and haven't listened to production audio this month, start there this week.

related

Keep reading