Clinical evaluation of a deep learning CBCT auto-segmentation software for prostate adaptive radiation therapy

Clin Transl Radiat Oncol 2024 Deep Learning 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Challenge of Prostate Radiation Therapy

Prostate cancer radiation therapy requires precisely targeting the prostate gland while protecting surrounding healthy organs. One major obstacle is geometrical uncertainty: the prostate and nearby structures such as the bladder and rectum can shift position from day to day and even during a single treatment session.

These movements mean that the radiation plan created at the start of treatment may not perfectly match the patient's anatomy on any given day. Dosimetric accuracy -- how closely the radiation dose matches the intended plan -- can be compromised over a course of treatment lasting several weeks.

Adaptive radiotherapy (ART) was developed to address this problem. ART involves regularly re-evaluating and adjusting the treatment plan to account for anatomical changes, and it requires fast, accurate identification of both the tumor target and surrounding organs on daily imaging scans.

The most practical imaging tool for this daily monitoring is the cone-beam computed tomography (CBCT) scanner built into the treatment machine. However, manually tracing the outlines of organs on CBCT images is slow and labor-intensive, limiting the practicality of adaptive approaches in busy clinical settings.

TL;DR: Daily shifts in prostate anatomy during radiation therapy demand adaptive re-planning, but manually outlining organs on imaging scans is too time-consuming for routine clinical use.
Page 2
Why Deep Learning for CBCT Segmentation?

Traditionally, a technique called deformable image registration (DIR) was used to transfer organ outlines drawn on the original planning CT scan onto daily CBCT images. While this works reasonably well for small anatomical changes, it can fail when the bladder or rectum is dramatically more or less full than at the time of the planning scan.

Deep learning auto-segmentation has become widely used for organ outlining on standard CT images, offering speed and accuracy comparable to expert clinicians. Extending this approach to CBCT could solve the limitations of DIR-based methods.

However, CBCT images have lower image quality than standard CT scans: they are noisier, have more artifacts, and show less contrast between soft tissues. These limitations challenge deep learning algorithms, which must learn to recognize organ boundaries that are less clearly defined.

An additional problem is that CBCT images are rarely labeled with organ outlines in routine clinical practice, which means there are far fewer training examples available for building algorithms compared to standard CT. Despite these hurdles, recent studies have shown promising results, and some radiation therapy machine manufacturers have begun integrating deep learning CBCT tools into their systems.

TL;DR: Deep learning offers a faster alternative to traditional image registration for daily organ outlining, but CBCT image quality and limited training data present real technical challenges.
Pages 2-3
Study Design and Data Collection

This study evaluated a commercial deep learning auto-segmentation software (referred to as DL) for its ability to accurately outline organs on CBCT images of prostate cancer patients. Ten patients who received definitive radiation therapy to the prostate gland and seminal vesicles were enrolled.

Patients received a moderately hypofractionated treatment schedule: 70 Gy to the prostate and 63 Gy to the seminal vesicles delivered in 28 fractions using a simultaneous integrated boost technique. CBCTs were acquired on a TrueBeam linear accelerator; one CBCT per patient (from fraction 3) was selected for analysis.

Four expert radiation oncologists independently drew outlines -- called contours -- for five structures on each CBCT: the prostate gland, seminal vesicles, femoral heads, bladder, and rectum. These four sets of contours were then combined using the STAPLE algorithm to create a consensus contour (CC), which served as the gold standard reference.

The same CBCTs were then processed by the DL software using two different models: one trained specifically on CBCT images (DL-CBCT) and one trained on standard CT images (DL-CT). This comparison was designed to test whether training on the target imaging modality matters.

TL;DR: Ten prostate cancer patients had one CBCT each contoured by four expert clinicians and by two deep learning models to test the software's accuracy against a clinical gold standard.
Page 3
Measuring Accuracy and Clinical Relevance

Three quantitative metrics were used to compare the deep learning contours against the consensus contour. The Dice similarity coefficient (DSC) measures spatial overlap between two contours on a scale from 0 (no overlap) to 1 (perfect overlap). Higher DSC values indicate better agreement.

The centre of mass (COM) shift measures how far the center of the deep learning contour is displaced from the center of the consensus contour, in millimeters, across three spatial directions. Smaller COM shifts indicate that the outlines are positioned in the same location.

The volume relative variation (VRV) measures the percentage difference in total contour size between the deep learning output and the consensus reference. Systematic over- or under-estimation of organ volume can affect the accuracy of radiation dose calculations.

Because no universally accepted tolerance thresholds exist for these metrics, the researchers used a novel benchmark: they compared the DL-CBCT performance against the natural variation between human experts (inter-observer variability, IOV). If the DL software performs no worse than the spread among human clinicians, it is considered clinically acceptable.

TL;DR: Accuracy was measured with three geometric metrics -- overlap, position, and volume -- and benchmarked against natural variation between human experts rather than arbitrary thresholds.
Pages 3-4
Performance Across Different Structures

The DL-CBCT software performed best on the femoral heads (hip bones), which are high-contrast bony structures easily visible on CBCT. Mean DSC was 0.96 for both DL-CBCT and DL-CT, with a center of mass shift below 2 mm -- results comparable to those seen with CT auto-segmentation.

For the bladder, DL-CBCT achieved a mean DSC of 0.90 and a center of mass shift of 2.2 mm. This good performance is explained by the bladder's relatively high contrast compared to surrounding tissue, although its interface with the prostate creates some difficulty.

The rectum showed a mean DSC of 0.86 for DL-CBCT, slightly lower than the bladder due to its irregular shape and variable filling. One consistent pattern noted by expert reviewers was that the deep learning contours tended to stop at the sigmoid colon cranially -- a systematic anatomical boundary choice that may differ from some institutional conventions.

The prostate gland achieved a DSC of 0.83 for DL-CBCT, compared to only 0.74 for DL-CT -- a meaningful difference that demonstrates the value of CBCT-specific training. For the seminal vesicles, DL-CBCT scored 0.70 while DL-CT dropped to 0.59, reflecting how challenging these small, low-contrast structures are to delineate on CBCT.

TL;DR: Accuracy ranged from excellent (DSC 0.96) for bony landmarks to moderate (DSC 0.70) for seminal vesicles, with the CBCT-trained model consistently outperforming the CT-trained model.
Page 4
Comparison with Human Expert Variability

The key clinical question was not just whether DL-CBCT is accurate in absolute terms, but whether it performs within the range of disagreement between human experts. For most structures and metrics, no statistically significant difference was found between DL-CBCT performance and the inter-observer variability among the four expert radiation oncologists.

Where statistically significant differences did exist -- notably for femoral heads -- the practical magnitude was small and clinically insignificant. For example, the median DSC was 0.94 for DL-CBCT comparisons versus 0.93 for expert-to-expert comparisons, a difference of 0.01 that would not affect treatment planning decisions.

An interesting finding was that in cases where DL-CBCT differed significantly from the inter-observer variability benchmark, the algorithm actually showed lower variability than the human experts -- meaning it was more consistent, not worse. For the rectum, the median volume variability was 15% for DL-CBCT versus 22% for the human experts.

This finding suggests that DL-CBCT can serve as a reliable starting point for clinical contouring. While expert review and correction remain necessary, especially for complex structures like the prostate and seminal vesicles, the AI-generated outlines provide a solid foundation that is well within the range of normal human clinical practice.

TL;DR: For nearly all structures, the deep learning software performed within the natural variation between human experts, and in some cases showed greater consistency than clinicians.
Pages 4-5
Clinical Implications for Adaptive Radiotherapy

The central finding -- that DL-CBCT performs comparably to human experts -- supports its integration into clinical adaptive radiotherapy workflows. Rather than replacing expert clinicians, the software can reduce the manual workload by providing an accurate starting point that requires only review and minor corrections.

The results also provide clear evidence that algorithms should be trained on the same imaging modality they will be used on. Training-modality matching consistently produced better results, particularly for soft tissue structures like the prostate and seminal vesicles where image contrast is limited on CBCT.

For online adaptive radiotherapy -- where the treatment plan is adjusted before each fraction while the patient is on the treatment table -- speed is critical. Automated segmentation eliminates a major bottleneck in this workflow. For offline ART, where re-planning happens between sessions, the time savings also allow more frequent plan updates.

The researchers note that human expert review remains necessary because deep learning algorithms can make specific systematic errors, such as the rectum boundary decisions noted in this study, and because clinical target volumes for the prostate may incorporate anatomical knowledge that goes beyond simple organ boundaries visible on imaging.

TL;DR: Deep learning CBCT segmentation can streamline adaptive radiotherapy workflows by providing clinically acceptable starting contours that require only expert review rather than full manual outlining.
Page 5
Conclusions and Future Directions

This study represents the first clinical evaluation of a deep learning auto-segmentation software specifically trained on CBCT images for prostate adaptive radiotherapy. The software achieved accuracy in line with consensus contours and within the inter-observer variability of expert clinicians, supporting its suitability for clinical use.

Several limitations should be acknowledged. The study involved only 10 patients, making this a preliminary evaluation rather than a definitive validation. The software evaluated was a beta version, and the final clinical release may incorporate further optimizations that could improve performance.

For integration into online adaptive radiotherapy, the auto-segmentation software would need to interface with dose calculation systems. This creates additional technical challenges related to CBCT image quality for dose calculation, including issues with electron density assignment and field-of-view limitations.

The authors conclude that the results are encouraging for the broader adoption of deep learning CBCT segmentation in clinical practice. Such tools are seen as an important step toward fully realized online adaptive radiotherapy, simplifying and speeding up the segmentation workflow that is a prerequisite for accurate daily dose evaluation and plan adaptation.

TL;DR: Despite being tested on only 10 patients with a beta software version, results support clinical adoption of CBCT deep learning segmentation as a key enabler of adaptive prostate radiotherapy.
Citation: Open Access, . Available at: PMC11176659.