Writing

An Online Automatic Corona Diagnose System Based on Chest X-ray Images

In short: I trained four networks on chest X-rays to detect COVID-19, picked the best one, and put it online. VGG16 won with 98.92% accuracy and it now runs behind an API on Google Cloud. 98.92% VGG16…

published
read time
6 min
words
1,223
lang
en

In short: I trained four networks on chest X-rays to detect COVID-19, picked the best one, and put it online. VGG16 won with 98.92% accuracy and it now runs behind an API on Google Cloud.

98.92%
VGG16 accuracy, the best of the four
4
architectures compared on the same data
1,400
X-ray images, split 80:20 into 1,120 train and 280 test

Why X-ray and not CT

SARS-CoV-2 spread across the world fast [1]. It infects the lungs and causes pneumonia in most patients. RT-PCR is reliable [2], but it takes time and some kits are not accurate enough. The common symptoms are fever and cough, plus shortness of breath, headache and fatigue [3].

Imaging helps. CT and X-ray are both used, but CT is not available in most small cities and it costs more. So this study is about X-ray, and about closing the gap between taking the image and getting the answer. The goal was an online system that reports lung engagement with the disease, patient status, and therapeutic guidelines, and takes some pressure off radiologists.

Other groups had gone at this too. Apostolopoulos and Mpesiana [4] compared five CNNs across normal, pneumonia and COVID-19 lungs, and found MobileNet v2 effective. Hemdan et al. [5] built COVIDX-Net, where VGG19 and DenseNet came out on top. Narin et al. [6] tested ResNet50, InceptionV3 and Inception-ResNetV2, and ResNet50 gave 98% accuracy.


How the system was built

Complete flow diagram of the study
Figure 1. Complete flow diagram of study
  1. Build the dataset400 confirmed positive COVID-19 subjects, from the public collection by Cohen et al. [8] plus images from hospitals in Ardabil province, Iran. 1,000 healthy lung X-rays came from the RSNA pneumonia detection challenge [12].
  2. Fine-tune four networksVGG16, VGG19, InceptionV3 and ResNet50, each with a fully connected layer added on top of the pre-trained model.
  3. Evaluate on the same termsSame epochs, same data, 80:20 split. Inputs at 244x244, except InceptionV3 at 299x299.
  4. Deploy the winnerThe best model goes onto Google Cloud Platform behind a Python API. Images and results move as JSON.

The models

VGG16 and VGG19 come from Simonyan and Zisserman [9]. They stack convolutional layers on top of each other. VGG16 has 138.4 million parameters, VGG19 has 143.7 million. Training networks that size on a small COVID-19 dataset is not efficient, so fine-tuning is what makes them usable here.

ResNet50 came from He et al. [10] in 2015, with 25.6 million parameters and 152 layers, eight times deeper than the VGG networks. Inception-V3 came from Szegedy et al. [11] the same year, around 23 million parameters, and hit 5.6% top-5 error for single frame evaluation on the ILSVRC 2012 challenge.

The API

The backbone sits on Google Cloud Platform, which gives serverless compute, so the model is cheap to create and run. Front-end and back-end talk in JSON. Input images are not stored, which avoids a storage problem and a privacy one. When an encoded image arrives, the Python core decodes it, preprocesses it, and analyses it. Three things come back to the website: lung engagement with the disease, patient status, and therapeutic guidelines. The first one is the real output. The other two follow from it.

TipPutting the model behind an endpoint instead of shipping it means every user is on the latest version the moment it is retrained. That was the point of doing it online.

Results

Everything ran in Python 3.7 on a laptop with an Intel Core i7-9750H, an Nvidia GTX 1650 2GB, and 24GB of RAM. Accuracy, sensitivity (recall) and specificity all come out of the confusion matrix: true positive, true negative, false negative, false positive.

Equations 1 to 3 for accuracy, sensitivity and specificity

VGG16 came out ahead, with the highest sensitivity and specificity as well as the highest accuracy. VGG19 was almost identical. The other two were not usable.

NetworkAccuracy (%)Sensivity (%)Specifity (%)
VGG1698.9296.25100
VGG1998.9097.5099.50
InceptionV371.791.25100
ResNet5028.271000.00

Table 1. Final results of networks

CarefulLook at the two failures in the confusion matrix, not just the accuracy column. InceptionV3 caught 1 positive out of 80. ResNet50 caught all 80 positives by calling every single image positive. Both are useless, and one of them still reports 71.79% accuracy.
NetworkTrue PositiveFalse NegativeFalse PositiveTrue Negative
VGG167730200
VGG197821199
InceptionV31790200
ResNet508002000

Table 2. Confusion matrix of CNN models

Here are the training histories. Each plot has training accuracy, validation accuracy, training loss and validation loss.

VGG19 training history plot
Figure 2. VGG19 training history plot
VGG16 training history plot
Figure 3. VGG16 training history plot
ResNet50 training history plot
Figure 4. ResNet50 training history plot
InceptionV3 training history plot
Figure 5. InceptionV3 training history plot

The ROC curve is below. The closer the curve hugs the true-positive border and then the top border of the ROC space, the more accurate the test.

Receiver operating characteristic curve of the four networks
Figure 6. The receiver operating characteristic curve of networks

Precision and F1-score are two more ways to judge the networks:

Equations 4 and 5 for precision and F1-score
Precision and F1-score of the CNN models
Table 3. Precision and F1-score of CNN models

Conclusion

Combining image processing and machine learning gives you convolutional networks that already do real work, in self-driving cars and elsewhere. Medicine is the obvious next place. The SARS-CoV-2 outbreak made the gap visible: as hospitals filled up and demand for CT and X-ray rose, the time to a confirmed diagnosis became the thing that mattered.

This study evaluated four CNNs, took the best one, and deployed it on GCP as an online diagnosing system. VGG16 outperformed the rest. The platform is free and gets updated regularly.

Acknowledgments

Thanks to everyone working in hospitals and caring for patients. This work would not have been possible without help from Ardabil University of Medical Sciences and Ardabil Science and Technology Park, Ardabil, Iran.

References

  1. Cohen, J., & Normile, D. (2020). New SARS-like virus in China triggers alarm, Science, vol. 367, no. 6475, pp. 234-235, 2020.
  2. Fang, Y., Zhang, H., Xie, J., Lin, M., Ying, L., Pang, P., & Ji, W. (2020). Sensitivity of chest CT for COVID-19: comparison to RT-PCR. Radiology, 200432.
  3. Wang, W., Tang, J., & Wei, F. (2020). Updated understanding of the outbreak of 2019 novel coronavirus (2019-nCoV) in Wuhan, China. Journal of medical virology, 92(4), 441-447.
  4. Apostolopoulos, I. D., & Mpesiana, T. A. (2020). Covid-19: automatic detection from x-ray images utilizing transfer learning with convolutional neural networks. Physical and Engineering Sciences in Medicine, 1.
  5. Hemdan, E. E. D., Shouman, M. A., & Karar, M. E. (2020). Covidx-net: A framework of deep learning classifiers to diagnose covid-19 in x-ray images. arXiv preprint arXiv:2003.11055
  6. Narin, A., Kaya, C., & Pamuk, Z. (2020). Automatic detection of coronavirus disease (covid-19) using x-ray images and deep convolutional neural networks. arXiv preprint arXiv:2003.10849
  7. Online COVID-19 Diagnose System.
  8. Cohen, J. P., Morrison, P., & Dao, L. (2020). COVID-19 image data collection. arXiv preprint arXiv:2003.11597.
  9. Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556
  10. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778)
  11. Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., & Wojna, Z. (2016). Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2818-2826)
  12. Stein, A., (2018) Pneumonia Dataset Annotation Methods. RSNA Pneumonia Detection Challenge Discussion. kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723

related

Keep reading