The Clinical Need Accurate HCC segmentation in MRI is essential for radiotherapy planning (defining treatment margins), surgical planning (guiding resection boundaries), and post-treatment monitoring (objectively assessing response). Manual segmentation is labor-intensive, variable between readers, and a bottleneck in both clinical practice and research workflows.
Why MRI Specifically? MRI offers superior soft tissue contrast and avoids ionizing radiation compared to CT, making it the preferred modality for HCC characterization - particularly in cirrhotic livers where gadolinium-enhanced hepatobiliary phase imaging is diagnostic. Deep learning for MRI-specific HCC segmentation requires a dedicated evidence synthesis.
The Review Scope From 2,462 records searched across PubMed, Scopus, Web of Science, and Cochrane Library, only 13 peer-reviewed studies met inclusion criteria (deep learning for HCC segmentation in MRI with quantitative metrics). This narrow evidence base highlights how early this field still is.
Key Finding Dice similarity coefficients (DSC) ranged from 0.61 to 0.954 across the 13 studies - a wide performance range reflecting the diversity of architectures, MRI protocols, dataset sizes, and annotation quality. U-Net-based models dominated with median DSC 0.83, while hybrid CNN-transformer models showed higher median DSC of 0.86.
Search Strategy A systematic literature search from January 2015 (coinciding with the publication of U-Net, the foundational architecture for medical image segmentation) across four databases identified 2,462 initial records. After deduplication and screening, 57 articles underwent full-text evaluation, with 13 meeting all criteria.
Inclusion Standards Studies required: deep learning models for HCC segmentation specifically in MRI; quantitative performance metrics (DSC, IoU, or sensitivity); human subjects with confirmed HCC; peer-reviewed original research; and MRI as the primary imaging modality. This strict focus on MRI-only, segmentation-only, DL-only studies explains the small final count.
Quality Assessment The QUADAS-2 tool was used to assess risk of bias across four domains: patient selection, index test, reference standard, and flow and timing. Most studies showed low bias in index test (11/13) and flow and timing (13/13) domains, but high risk of bias in patient selection was noted in 8 of 13 studies due to non-consecutive enrollment or unclear inclusion criteria.
Performance Synthesis Because of substantial heterogeneity in architectures, MRI protocols, and dataset characteristics, narrative synthesis with grouped analysis by architecture type (U-Net, transformer, hybrid) and MRI sequence was performed rather than pooled meta-analysis.
U-Net Dominance U-Net and its variants (UNet++, ResUNet, nnU-Net, 3D U-Net) appeared in 7 of 13 studies and achieved a median DSC of 0.83 (range 0.74-0.954). The highest-performing study used AS-Net, a 3D U-Net applied across all BCLC stages, achieving DSC 0.954 +/- 0.018 with a robust multi-expert annotation protocol.
Transformers and Hybrid Models Two transformer-based studies achieved median DSC of 0.79 (range 0.772-0.829), while two hybrid CNN-transformer models showed the highest median DSC of 0.86 (range 0.83-0.89). The attention mechanisms in transformers may capture long-range spatial dependencies across large HCC tumors better than purely local convolution-based approaches.
Impact of Annotation Quality Studies with robust multi-expert annotation protocols (such as multiple senior radiologist reviews) consistently achieved higher DSC scores. Studies with unclear inter-observer reliability in ground truth annotation showed lower and more variable DSC scores, confirming that annotation quality is a key driver of reported performance.
Small Lesion Challenge Multiple studies reported substantially lower performance for small HCC lesions (under 2 cm), with some architectures showing F1-scores of 0.59 for small lesions compared to 0.76 for human raters - precisely the lesion size range where deep learning assistance is most needed clinically for early HCC detection.
Clinical Application Domains Across the 13 studies, deep learning HCC segmentation in MRI was applied to diagnosis support (confirming HCC boundaries for LI-RADS classification), treatment planning (defining margins for stereotactic body radiotherapy and surgical resection), and risk assessment (measuring tumor volume for staging and response evaluation).
Dataset Size Limitation Study datasets ranged from 19 to 602 patients - far smaller than the thousands of annotated scans needed for robust deep learning in medical imaging. Small dataset size is the single most consistently cited limitation across all 13 studies, constraining model complexity and generalizability.
MRI Protocol Variability Different scanner field strengths (1.5T vs. 3T), MRI sequences (T1, T2, dynamic contrast-enhanced), gadolinium agents, and hepatobiliary phase timing across institutions create substantial domain shift that causes models trained at one institution to underperform when deployed at another.
Lesion Heterogeneity HCC demonstrates marked intratumoral heterogeneity including areas of necrosis, hemorrhage, and satellite nodules that create ambiguous boundaries on MRI. This biological variability means that even expert radiologists disagree on HCC boundaries, contributing to the wide performance range observed across studies.
Radiotherapy Planning Stereotactic body radiotherapy (SBRT) for unresectable HCC requires precise gross tumor volume delineation as the foundation for treatment planning. Manual MRI segmentation for SBRT planning can take 30-60 minutes per patient; automated deep learning segmentation could reduce this to seconds, enabling broader adoption of liver SBRT.
Transplant Staging Accuracy Milan criteria and University of California San Francisco (UCSF) criteria for HCC transplant eligibility depend on accurate tumor number and maximum diameter measurement. AI segmentation providing precise, reproducible measurements could reduce staging errors that lead to inappropriate transplant exclusions or inclusions.
MRI over CT Advantage For patients with advanced cirrhosis where contrast-enhanced CT carries higher renal contrast risk, MRI-based AI segmentation provides an alternative imaging pathway for tumor burden assessment without iodinated contrast - important for the chronic kidney disease that frequently accompanies end-stage liver disease.
Longitudinal Monitoring Gadoxetate disodium (Eovist/Primovist) hepatobiliary MRI is increasingly used for HCC surveillance in high-risk populations. Automated AI segmentation applied to serial MRI exams could enable consistent, standardized interval comparison of tumor size and count without reader-dependent variability.
Multi-Center Datasets Urgently Needed The most impactful next step identified by this review is the establishment of large, annotated multi-center MRI datasets for HCC segmentation. Collaborative efforts similar to TCIA (The Cancer Imaging Archive) for CT are needed to provide the training and validation data required for generalizable MRI models.
Standardized MRI Protocols Clinical deployment requires models that perform consistently across scanner vendors and field strengths. Standardized acquisition protocols for research HCC segmentation - analogous to the LI-RADS v2018 imaging technical requirements - would reduce domain shift and enable fairer cross-study performance comparisons.
Hybrid Human-AI Approaches Given the remaining performance gap, particularly for small lesions, the near-term clinical deployment model should be hybrid: AI generates initial segmentations that radiologists rapidly review and correct, rather than fully autonomous AI segmentation. This maintains clinical accountability while capturing the efficiency gains of automation.
Prospective Clinical Validation All 13 reviewed studies were retrospective. Prospective studies embedding AI HCC segmentation into actual clinical radiology workflows are needed to quantify real-world time savings, identify failure modes not captured in retrospective testing, and demonstrate impact on clinical decision-making.