Respiratory sound classification, end to end
In short: I took a public set of lung sound recordings all the way from raw audio to a web demo anyone can upload a file to. The model was the smallest part. Most of the work was deciding what to…
- published
- read time
- 4 min
- words
- 854
- lang
- en
- filed under
- Engineering
In short: I took a public set of lung sound recordings all the way from raw audio to a web demo anyone can upload a file to. The model was the smallest part. Most of the work was deciding what to predict, splitting the data honestly, and making every run comparable.
I built it in a short amount of time, so I treated it as a rehearsal of how I would approach a classification problem like this at full size: try several options at every stage, log everything, and let the comparison pick the winner.
Decide what to predict before anything else
The data is the ICBHI 2017 respiratory sound database: recordings from a stethoscope, a text file of annotations per recording, and a diagnosis per patient. The first thing exploration showed was that it is highly imbalanced. Some diagnoses, asthma for example, have only a handful of recordings.
You cannot train a classifier for a class you have five examples of, so I did not try. I framed two tasks instead:
- Binary: normal or abnormal.
- Multi-class: normal, chronic respiratory diseases, and respiratory infections.
Grouping diagnoses into broader buckets is a product decision as much as a modelling one. It trades detail for classes that have enough data to learn from. That trade should be written down, because someone will later ask why asthma is not on the list.
Three ways to turn sound into numbers
A lung recording is long and mostly quiet, with short events like crackles and wheezes on top of breathing. I compared three inputs: MFCCs, a log-mel spectrogram, and MFCCs with augmented features. The log-mel keeps more detail across frequency. MFCCs compress it into a few smooth coefficients per frame, which is smaller and often enough.
The feature code is short. Something like this, with the real version in utils/audioprocessing.py:
import librosa
import numpy as np
def features(path, kind="mfcc", sr=22050):
y, sr = librosa.load(path, sr=sr)
if kind == "mfcc":
return librosa.feature.mfcc(y=y, sr=sr, n_mfcc=40)
if kind == "log_mel":
mel = librosa.feature.melspectrogram(y=y, sr=sr, n_mels=128)
return librosa.power_to_db(mel, ref=np.max)
raise ValueError(kind)
Training: one script, every combination
Train.py takes the task and the input type as arguments, and can run all six combinations in one go. Each run does the same steps in the same order:
- Preprocess: filtering, resampling, feature extraction.
- Split into training, validation and test sets.
- Oversample the minority classes with SMOTE, on the training set only.
- Tune hyperparameters with Optuna.
- Log parameters and metrics to MLflow, and save the model.
Step 3 is where the easy mistake lives. If you oversample before splitting, synthetic copies of a test example end up in training, and your test score is a lie. Split first, then oversample.
python Train.py --debug runs the whole pipeline on random data. It catches broken paths and shape errors in seconds, before you spend an hour extracting features.Evaluation: what accuracy hides
The test sets are saved as .npy files that no training run ever touched, and TestModels.py matches each one to its model and scores it. On data this imbalanced, accuracy alone tells you very little. Here is why, on an illustrative split where nine in ten recordings are abnormal. These are the scores of models that learned nothing:
| Model that learned nothing | Accuracy | Recall, normal | Recall, abnormal | Macro F1 |
|---|---|---|---|---|
| Always says "abnormal" | 0.90 | 0.00 | 1.00 | 0.47 |
| Guesses by class share | 0.82 | 0.10 | 0.90 | 0.50 |
| Always says "normal" | 0.10 | 1.00 | 0.00 | 0.09 |
A model that always says "abnormal" scores 0.90 accuracy and is useless. So I read per-class recall and macro F1 first, and I treat these rows as the floor a real model has to clear. The real scores for each model are on the model performance page of the demo.
One thing I would tighten with more time: splitting by patient rather than by recording. The same patient appears in several recordings, and a model can learn the patient instead of the disease.
Shipping it as a demo
The front end is a Streamlit app, app.py, with pages for the data exploration, the model comparison, and an upload box that runs inference. A GitHub Actions workflow handles a limited CI/CD path. Run locally, Prometheus collects the app's metrics and Grafana draws them.
None of that changes the model. It changes whether anyone else can check it. A model that only exists in a notebook cannot be questioned by a clinician, and the questions are the point.
Run it yourself
Record or find a .wav of breathing and run it through the pipeline. If you build your own, give the training script a debug flag that runs the whole path on random data first, then try one real configuration, binary labels and MFCCs, before the full grid. Before you look at the accuracy, compute the "always abnormal" row for your own split. That is the number to beat.
related