Convolutional Neural Networks for Automated Classification of Prostate Multiparametric Magnetic Resonance Imaging Based on Image Quality

J Magn Reson Imaging 2022 Deep Learning 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Prostate MRI Quality Control Matters

Prostate multiparametric MRI (mpMRI) has become central to the early detection of clinically significant prostate cancer, enabling targeted biopsy of suspicious lesions and reducing the number of unnecessary procedures. European Association of Urology guidelines now recommend MRI as the first diagnostic step for men with suspected prostate cancer, reflecting the strong evidence base for the MRI-guided pathway.

Despite the strong evidence supporting mpMRI, its diagnostic performance varies considerably between centers. Factors that affect quality include the MRI equipment used, the imaging protocol, proper patient preparation, and the radiologist's experience. A scan of insufficient quality can miss cancers, create false positives, or yield uninterpretable results, undermining the entire value of the MRI pathway.

The PI-RADS (Prostate Imaging Reporting and Data System) v2.1 guidelines establish minimum technical standards for prostate MRI acquisition, but adherence to these standards does not guarantee diagnostic quality. Patient-related factors such as movement during scanning and the presence of air in the rectum can degrade image quality regardless of protocol compliance.

Current quality assessment relies on visual review by the radiologist after the scan is complete, which is not always feasible in busy clinical settings and may not be possible in time to allow a repeat acquisition while the patient is still in the scanner. An automated, real-time quality classification system running at the MRI workstation could alert technologists immediately, enabling corrective action before the patient leaves.

TL;DR: Inconsistent prostate MRI quality across centers threatens the diagnostic value of the MRI pathway, motivating the development of automated real-time quality classification to enable immediate corrective action during scanning.
Pages 2-3
Study Design: 316 Patients, Four MRI Sequences, 28 CNN Architectures

This retrospective study included 316 prostate mpMRI scans from 312 men (median age 67) acquired on a single 3 Tesla MRI scanner between January and July 2020. Four sequence types were evaluated: T2-weighted imaging (T2WI), diffusion-weighted imaging (DWI) at b = 1500 sec/mm2, the derived apparent diffusion coefficient (ADC) maps, and dynamic contrast-enhanced (DCE) perfusion imaging.

Three genitourinary radiologists with 21, 12, and 5 years of experience independently reviewed each sequence and assigned a binary quality label: Q0 (low quality, insufficient for diagnosis) or Q1 (high quality, sufficient for diagnosis). Quality was assessed based on field of view adequacy, spatial resolution, signal-to-noise ratio, motion artifacts, magnetic susceptibility artifacts, rectal gas, and enhancement pattern on DCE. When readers disagreed, the majority label was used as the reference standard.

Twenty-eight convolutional neural network (CNN) architectures were trained and compared, including AlexNet, VGG variants, ResNet variants, DenseNet variants, GoogLeNet, ShuffleNet, MobileNet, ResNeXt, Wide ResNet, and MNASNet. All networks were initialized using transfer learning from weights pretrained on the ImageNet dataset, then fine-tuned on the prostate MRI quality classification task.

Training used a 10-fold cross-validation strategy with a 70/20/10 train/validation/test split. To prevent data leakage, all slices from a given patient were kept in the same fold. Data augmentation techniques including random rotation, vertical flipping, random shift, and elastic deformation were applied during training to increase effective dataset size and reduce overfitting. Majority vote aggregation was used to convert per-slice predictions into a single per-sequence label.

TL;DR: Twenty-eight CNN architectures were trained and compared on 316 prostate mpMRI scans labeled by expert radiologists, using transfer learning and 10-fold cross-validation across four sequence types to identify low-quality scans automatically.
Pages 2-3
CNN Architecture Diversity: From AlexNet to ShuffleNet

The 28 CNN architectures tested represent several distinct design philosophies. Spatial exploitation CNNs like AlexNet and VGG extract features through progressively larger receptive fields. Depth and multipath CNNs like ResNet use residual learning with identity mapping skip connections, allowing much deeper networks to be trained without the vanishing gradient problem.

DenseNet connects every layer to every other layer, maximizing cross-layer information flow. GoogLeNet pioneered the block concept with split-transform-merge operations, enabling efficient multi-scale feature extraction. Width-based CNNs like WideResNet and ResNeXt expand the number of feature map channels rather than depth, providing complementary representational power.

Lightweight architectures like MobileNet use depth-wise separable convolutions to drastically reduce computation, making them suitable for deployment on the modest hardware of an MRI workstation. ShuffleNet uses channel shuffling to enable efficient cross-group information flow in grouped convolutions, achieving strong accuracy with low computational cost.

All 28 architectures were modified only at their final layer, setting the output to two classes (Q0 and Q1). This approach leverages features learned from millions of natural images and redirects them toward the specific task of detecting image quality artifacts in medical MRI, a powerful form of domain adaptation via transfer learning that compensates for the relatively small size of the medical training dataset.

TL;DR: Twenty-eight CNN architectures spanning spatial exploitation, residual, dense, and lightweight design families were evaluated to identify which architecture families best detect image quality problems across the four prostate MRI sequence types.
Pages 4-5
Per-Slice and Per-Sequence Classification Performance

On the per-slice analysis, the best-performing CNN for each sequence achieved the following global accuracies: 89.95% for T2WI (VGG11), 79.83% for DWI (ResNet152), 76.64% for ADC (DenseNet161), and 96.62% for DCE (ShuffleNet v2-x1-0). DCE achieved the highest accuracy, while ADC was the most challenging sequence to classify.

When slice-level predictions were combined using a majority vote aggregation function to classify entire sequences, accuracy improved dramatically: 100% was achieved for T2WI, DWI, and DCE sequences, and 92.31% overall for ADC (with Q0-specific accuracy of 83.33%). This near-perfect sequence-level performance reflects the fact that quality-degrading artifacts typically affect most slices within a sequence, so correct classification of the majority of slices reliably determines the overall sequence quality.

The three best-performing architectures for each sequence did not differ significantly from each other (P > 0.05), but all significantly outperformed the remaining 25 architectures (P < 0.05). This suggests that architecture choice within the top tier matters less than avoiding the worst-performing architectures, which showed Q0-class accuracy as low as 14% for DWI. The choice of architecture is thus important primarily for avoiding poor classification of the rare but clinically critical low-quality cases.

Inter-reader agreement was almost perfect for T2WI and DCE (Fleiss kappa 0.83 and 0.80) and substantial for DWI and ADC (kappa 0.77 and 0.75). The lower agreement on DWI and ADC reflects the greater technical complexity and subjective interpretation of artifacts on these sequences, which also explains why the CNN models performed slightly less well on DWI and ADC than on T2WI and DCE.

TL;DR: CNN classifiers achieved near-perfect accuracy (100%) for identifying low-quality T2WI, DWI, and DCE sequences at the per-sequence level, with ADC slightly lower at 92.31%, using majority vote aggregation across individual slice predictions.
Pages 6-7
Clinical Integration: Real-Time Quality Control at the MRI Scanner

The most important clinical application of automated quality classification is real-time integration at the MRI workstation, where the algorithm would immediately notify the MRI technologist when a sequence is classified as low quality. This creates an opportunity to retake the sequence before the patient leaves the scanner, avoiding delays and repeat appointments that currently occur when suboptimal scans are identified only during radiologist review.

Automated quality control is particularly valuable because MRI technologists, who are responsible for image acquisition, may lack the clinical expertise to recognize the diagnostic implications of certain artifacts. For example, susceptibility artifacts from rectal gas on DWI may not appear severe to a technologist but significantly impair the detection of lesions in the posterior peripheral zone where most prostate cancers arise.

The clinical importance of detecting low-quality scans differs by sequence. DWI is the most critical sequence for peripheral zone assessment in prostate cancer detection and is the most susceptible to artifacts, making automated quality flagging most valuable for DWI. DCE, by contrast, plays a more limited role in PI-RADS scoring and is only used to upgrade selected PI-RADS 3 lesions, so a low-quality DCE has less clinical impact, and a repeat acquisition is often not feasible due to contrast timing constraints.

The ideal deployment scenario would be CNN models embedded in the MRI scanner's reconstruction software by the manufacturer, enabling on-the-fly quality assessment after each sequence is acquired. This would require vendor-specific training for each scanner model and protocol, since image characteristics and artifact patterns vary between equipment manufacturers. The present study establishes the proof of concept; vendor-specific deployment would require additional development and regulatory approval.

TL;DR: Automated CNN quality classification integrated at the MRI workstation could enable real-time alerts to technologists during scanning, preventing diagnostic failures from suboptimal DWI sequences that are most critical for prostate cancer detection.
Pages 7, 10
Limitations and the Path Toward Generalizable Quality Control

The primary limitation of this study is that all CNN models were trained on images from a single MRI scanner and acquisition protocol. It is currently unknown whether the trained classifiers would generalize to images acquired on different scanner models, field strengths, or protocols from other institutions. Multi-scanner, multi-site validation is required before deployment as a generalizable quality control tool.

The quality labels assigned by radiologists, which serve as the training reference, are inherently subjective. While inter-reader agreement was substantial to almost perfect across sequences, the binary Q0/Q1 classification does not capture the continuous spectrum of image quality. A finer grading system aligned with structured scoring tools like PI-QUAL could provide more nuanced feedback and may better support clinical decision-making.

The classifiers identify low-quality sequences but do not pinpoint the specific cause of the quality problem. A more advanced system capable of identifying whether the issue was caused by motion, magnetic susceptibility artifacts, inadequate field of view, or insufficient signal-to-noise ratio would be more actionable, helping technologists make the specific correction needed rather than simply flagging a problem.

Despite these limitations, this study demonstrates that CNNs can perform automated prostate MRI quality classification with accuracy approaching and in some cases matching that of expert radiologists at the per-sequence level. This represents a meaningful step toward standardizing prostate MRI quality across diverse clinical settings, which is a prerequisite for realizing the full diagnostic potential of MRI-guided prostate cancer detection.

TL;DR: CNN-based quality classification of prostate mpMRI achieves near-perfect per-sequence accuracy but requires multi-scanner validation and refinement to identify specific artifact causes before it can serve as a generalizable clinical quality control tool.
Citation: Open Access, . Available at: PMC9291235.