MRI is the preferred imaging modality for detecting and localizing prostate cancer, guided by the PI-RADS (Prostate Imaging Reporting and Data System) scoring framework. PI-RADS assigns cancer suspicion scores to lesions based on which anatomical zone they occupy -- requiring accurate delineation of the peripheral zone (PZ), where most cancers originate, and the transition zone (TZ), where benign hyperplasia is common.
The PI-RADS scoring criteria differ by zone: diffusion-weighted imaging (DWI) governs scoring in the peripheral zone, while T2-weighted imaging governs the transition zone. Without knowing which zone a suspicious lesion lies in, a radiologist cannot correctly apply PI-RADS criteria -- making accurate zone boundaries a prerequisite for the entire cancer detection workflow.
Beyond cancer detection, zonal segmentation enables reproducible prostate volume measurement, PSA density calculation (PSA divided by prostate volume -- a key diagnostic metric), MRI-ultrasound fusion biopsy targeting, and radiation therapy planning. These diverse downstream uses make automated zone segmentation a foundational capability for modern prostate cancer management.
Manual zone segmentation by radiologists is extremely time-consuming and subject to significant inter- and intra-observer variability, particularly at the prostate apex and base where boundaries are least clear. The prostate has fuzzy boundaries, heterogeneous internal intensity, and similar appearance to adjacent tissues -- making automated segmentation technically challenging.
This systematic review followed PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, with the review protocol pre-registered on PROSPERO before data collection began. Pre-registration prevents outcome reporting bias by committing to analysis plans in advance.
Four databases (Medline, ScienceDirect, Embase, and Web of Science) were searched for studies proposing automated machine learning or deep learning methods for prostate zone segmentation on MRI. Two radiologists with extensive prostate MRI experience independently screened 458 articles after duplicate removal, with conflicts resolved by a third senior radiologist.
Articles were included only if they used fully automated methods with manual segmentation as ground truth and reported standard performance metrics. Studies were excluded if they used semi-automated methods, segmented only the whole prostate gland without zonal breakdown, or did not provide quantitative performance evaluation. This rigorous selection yielded 33 final articles spanning 2011 to 2021.
Quality assessment used the QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) framework supplemented by the CLAIM (Checklist for Artificial Intelligence in Medical Imaging) criteria. QUADAS-2 evaluates risk of bias across four domains: patient selection, index test, reference standard, and flow/timing -- providing a structured framework for identifying methodological weaknesses in diagnostic accuracy studies.
Before 2017, prostate zone segmentation relied on atlas-based registration (warping labeled reference images onto new cases), C-means clustering, and classical machine learning approaches. These methods required substantial manual feature engineering and worked best when prostate anatomy was relatively typical.
After 2017, deep learning with convolutional neural networks (CNNs) became dominant, appearing in 72% (24 of 33) of reviewed studies. The fundamental architecture shift was from hand-crafted features to learned representations: networks like U-Net, V-Net, and ResNet automatically learned which image features distinguish zone boundaries without explicit programming.
U-Net became the backbone of choice, with researchers applying numerous modifications: combining multiple U-Nets, adding attention mechanisms (squeeze-and-excitation blocks and feature pyramid attention), and introducing atrous (dilated) convolutions for larger receptive fields without increased computational cost. Each modification aimed to improve accuracy on the specific challenge of prostate zone boundaries.
Performance varied by zone: nearly all studies found lower Dice Similarity Coefficients for the peripheral zone than the whole gland or transition zone. The peripheral zone's complex shape, thin anterior dimension, and overlap with surrounding tissues make its boundaries more ambiguous than the more centrally located transition zone. Typical best-performing DSC values across studies ranged from 0.70 to 0.78 for PZ versus 0.85 to 0.93 for the transition zone.
The review found 18 different types of zonal anatomy terminology across the 33 studies -- a strikingly high number given that standard prostate MRI anatomy recognizes exactly four zones: peripheral zone, transition zone, central zone, and anterior fibro-muscular stroma (AFMS). This terminological chaos makes comparing results across studies effectively impossible.
The most common error was inappropriate use of the term "central gland" (CG) -- a vague umbrella term not standardized in PI-RADS v2.1 -- to refer to regions that some authors defined as including the central zone plus transition zone plus AFMS, while others excluded the AFMS, and yet others conflated the central zone with the transition zone entirely. Two studies explicitly misused "central zone" to mean "central gland."
Only 8 of 33 articles provided precise terminology and a clearly documented segmentation protocol. The remaining studies had ambiguous or incorrect anatomical definitions that would make it impossible to reproduce their methods or meaningfully interpret their reported performance numbers -- since a DSC for "central gland" in one study may refer to a fundamentally different anatomical region than the same term in another study.
This terminology problem is not merely academic: prostate cancer risk differs substantially between zones, and PI-RADS scoring criteria depend on knowing whether a lesion is in the peripheral zone or transition zone. An automated segmentation model trained on ambiguously defined zones could misclassify lesions in a way that directly harms clinical decisions.
Using QUADAS-2, only 2 of 33 studies were judged to have low risk of bias in all four domains. Only one-quarter of studies (8 of 33) had low risk of bias for patient selection. Only one-third (10 of 33) had low risk for reference standard quality. The overall picture is a literature with widespread methodological weaknesses that undermine confidence in reported performance numbers.
Patient selection bias was pervasive: most databases lacked representativeness of patient variability including prostate volume range, heterogeneity of prostate tissue, and presence of both cancerous and benign hyperplasia cases. Models trained on homogeneous datasets -- such as only patients without prostate cancer -- produce systematically biased training labels because cancer changes prostate contour appearance.
Ground truth quality problems were common. Only 4 of 33 studies assessed inter-rater variability for their manual segmentations, and only 2 used blinded reading. Most studies trained on segmentations from a single reader without quantifying intra-reader consistency. Meyer et al. showed directly that training on single-reader segmentations introduces bias: models performed substantially better when evaluated by the same reader who created the training data than by a different expert.
Validation methodology was also a major weakness: none of the 33 studies used prospective data. Only 7 studies tested on both private and public datasets -- meaning most models were never evaluated on data from institutions other than where they were trained. This absence of external validation is a critical barrier to clinical deployment, since most segmentation models show significant performance degradation when applied to images from different scanners, field strengths, or acquisition protocols.
The three most commonly used public datasets were PROSTATEx (used for lesion classification challenges), NCI-ISBI 2013, and PROMISE12 (designed for the MICCAI prostate segmentation challenge). While these provide a shared benchmark, they have individual limitations: PROSTATEx uses 3.6 mm slice thickness that does not meet PI-RADS v2.1 quality requirements for T2-weighted images.
Scanner heterogeneity was inadequately addressed: fewer than half (14 of 33) of studies used data from more than one scanner vendor, and only 7 included both 1.5T and 3T MRI machines. Because MRI signal intensity is not standardized across vendors and field strengths, a model trained exclusively on one scanner type often performs poorly when tested on images from a different machine -- a failure mode that could prevent clinical deployment.
Most studies (24 of 33) used only T2-weighted imaging as the sole input sequence -- the appropriate choice since T2WI provides the clearest anatomical delineation for zone boundaries. A minority added ADC maps derived from diffusion-weighted imaging, which some studies found improved peripheral zone segmentation where the two modalities provide complementary boundary information.
Slice thickness compliance with PI-RADS v2.1 recommendations (3 mm or less for axial T2WI) was met in only 13 of 33 studies. Thicker slices create partial volume effects at zone boundaries and reduce the precision of segmentation ground truth -- yet this basic technical quality requirement was frequently overlooked even in published benchmark studies.
The review identifies the urgent need for a consensus segmentation protocol and standardized terminology for prostate zones, analogous to what already exists for organ-at-risk delineation in radiation therapy planning. Without agreed definitions of exactly which tissue regions each zone label should include, no valid cross-study comparison is possible.
Future database development should require multi-reader annotation with documented inter-rater variability, well-defined inclusion and exclusion criteria, representation of patient variability including a range of prostate volumes and pathologies (both cancer and benign hyperplasia), and adherence to PI-QUAL (Prostate Imaging Quality) standards to ensure image acquisition quality is sufficient for zone delineation.
External validation on data from multiple institutions is essential before any segmentation model can be considered ready for clinical use. The consistent finding across other AI imaging domains -- that models trained at one center degrade significantly at others -- makes it a near-certainty that unpublished performance drops exist for prostate segmentation models as well.
The adoption of structured reporting standards like the CLAIM (Checklist for Artificial Intelligence in Medical Imaging) checklist would help future studies provide sufficient technical detail for reproducibility and fair comparison. Currently, key information such as preprocessing steps, annotation tools, reader experience, and implementation details are inconsistently reported across the literature.
This systematic review reveals that while automated prostate zone segmentation using deep learning has made substantial technical progress since 2017, the published literature is not yet ready to support reliable clinical deployment. No studies have both sufficiently documented datasets and sufficient external validation to make their methods clinically trustworthy.
The field faces a compound problem: a lack of standardized terminology makes cross-study comparison impossible; biased training datasets with single-reader annotations and homogeneous patient populations inflate reported performance; and the near-universal absence of external validation leaves the real-world generalizability of these models completely unknown.
Deep learning methods -- primarily U-Net variants -- achieve strong performance on the transition zone (DSC 0.85-0.93) and whole gland, but consistently underperform on the peripheral zone (DSC 0.70-0.78), which is precisely where most clinically significant prostate cancers originate and where accurate zone segmentation matters most for PI-RADS scoring.
The development of high-quality, multi-institutional, multi-reader databases with standardized zone definitions and image quality criteria is an essential prerequisite for the next generation of clinical-grade prostate MRI segmentation tools -- tools that could ultimately reduce manual workload, improve PI-RADS consistency, and enable more reliable automated cancer detection.