Deep learning on automated performance metrics and clinical features to predict urinary continence recovery after robot-assisted radical prostatectomy

BJU Int 2019 Deep Learning 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Continence Recovery: A Critical Quality-of-Life Outcome After Prostate Surgery

Robot-assisted radical prostatectomy (RARP) is the most common surgical treatment for localized prostate cancer, with over 90% of radical prostatectomies in the United States now performed robotically. Despite technical advances, urinary incontinence remains one of the most impactful complications, with recovery taking anywhere from days to over a year after surgery.

Current methods for assessing surgical quality rely heavily on surgeon self-reporting and subjective mentorship evaluations, which are inconsistent and difficult to standardize across training programs. There is no validated objective framework for quantifying surgical technique or predicting patient-specific functional outcomes before or shortly after surgery.

Predicting which patients will recover continence quickly versus slowly has significant clinical value: it enables personalized patient counseling before surgery, allows early identification of patients who may need targeted rehabilitation, and provides an objective metric for comparing surgical technique across surgeons and institutions.

Existing predictive models for continence recovery rely on clinical features such as patient age, body mass index, prostate volume, and preoperative continence status -- but these models have limited accuracy because they do not capture the quality of the surgical procedure itself.

TL;DR: Urinary continence recovery after robotic prostate surgery is a major quality-of-life concern that current clinical models cannot reliably predict because they ignore the quality of surgical execution.
Pages 2-3
Automated Performance Metrics: Quantifying the Surgical Act

Automated performance metrics (APMs) are kinematic and motion-based measurements automatically extracted from the da Vinci robotic surgical system during an operation without any additional sensors or manual annotation. They include instrument motion economy metrics such as path length, angular range of motion, depth range, and speed for each of the surgeon's hand instruments.

APMs capture subtle differences in how surgeons move their instruments during specific operative steps. An expert surgeon performing vesico-urethral anastomosis -- the critical suturing step that reconnects the bladder to the urethra after prostate removal -- will typically show more efficient, economical motion compared to a less experienced surgeon performing the same task.

Prior work has validated APMs as correlates of surgical experience and skill in several procedures, but their relationship to patient outcomes rather than technical skill scores had not been thoroughly investigated. This study is among the first to directly link APMs to a clinically meaningful functional outcome -- urinary continence -- rather than to process metrics.

492 APMs were extracted per case across 23 defined surgical task segments, covering both instrument-specific metrics and combined bimanual metrics. Alongside APMs, 16 clinical features (patient demographics, preoperative status, and surgical case characteristics) were collected for each of 100 RARP cases performed by 8 surgeons at a single institution.

TL;DR: Automated performance metrics extracted from the da Vinci robotic system capture objective, continuous measurements of surgeon hand motion that may reflect technical quality beyond what clinical features alone can measure.
Pages 3-4
DeepSurv: A Deep Learning Approach to Time-to-Event Prediction

Continence recovery is a time-to-event outcome -- the clinically relevant question is not just whether a patient recovers continence, but how quickly. This is modeled as a survival analysis problem, where the event of interest is return to continence and the time variable is days after surgery.

DeepSurv is a deep neural network extension of the classical Cox proportional hazards model that can learn nonlinear relationships between high-dimensional input features and the hazard (risk of the event over time). Unlike traditional survival models, DeepSurv does not require the analyst to prespecify which features matter or how they interact.

Three models were compared: a clinical-only model using 16 clinical features, an APMs-only model using 492 APMs, and a combined model incorporating all features together. All three used the DeepSurv architecture with identical hyperparameters, allowing direct comparison of information contributed by each feature type.

Model performance was measured using concordance index (CI), which measures whether patients predicted to recover faster actually do recover faster (a value of 1.0 is perfect, 0.5 is random), and mean absolute error (MAE), which measures the average difference in days between the predicted and actual time to continence recovery.

TL;DR: DeepSurv, a deep learning survival model, was applied to predict time-to-continence recovery after robotic prostatectomy, using three parallel models to isolate the predictive contribution of APMs versus clinical features.
Pages 4-5
APMs Drive Predictive Performance

The combined DeepSurv model (APMs plus clinical features) achieved a concordance index of 0.60 and a mean absolute error of 85.9 days for predicting time to urinary continence recovery, compared to a CI of 0.55 for clinical features alone and 0.58 for APMs alone, indicating meaningful but modest improvement from each feature type.

Feature importance analysis using SHAP values revealed that all 10 top-ranked predictive features came from APMs -- no clinical feature appeared in the top 10. The most informative APMs were instrument motion metrics during two critical surgical steps: vesico-urethral anastomosis (reconnecting the bladder to the urethra) and prostatic apical dissection (removing the prostate apex near the urinary sphincter).

Among APM types, instrument path length and angular range of motion during anastomosis were consistently among the highest-ranked features, consistent with the interpretation that more economical, precise suturing at the anastomosis site better preserves the urinary sphincter mechanism.

In the full cohort of 100 patients, 79% of patients recovered continence with a median recovery time of 126 days after surgery, providing a realistic baseline for evaluating model predictions and confirming that the dataset was appropriately representative of contemporary RARP outcomes.

TL;DR: The combined DeepSurv model achieved a concordance index of 0.60, with APMs from vesico-urethral anastomosis and apical dissection dominating the top 10 predictive features while no clinical variable reached that threshold.
Pages 5-6
APM-Based Surgeon Grouping Predicts Better Early Continence Outcomes

To test whether APMs could distinguish surgeons whose patients recover continence more quickly, surgeons were grouped using k-means clustering on their APM profiles into Group 1 (more efficient motion patterns) and Group 2 (less efficient motion patterns). This grouping was compared against a traditional experience-based grouping by years of RARP experience.

In a historical validation cohort of 493 RARP cases, patients operated on by Group 1 (APM-efficient) surgeons achieved a 3-month continence rate of 47.5% compared to 36.7% for Group 2 surgeons (p = 0.034), and a 6-month continence rate of 68.3% versus 59.2% (p = 0.047). The difference was not statistically significant at 12 months, suggesting APM-related advantages are strongest in the early recovery period.

Critically, experience-based grouping (by years performing RARP) did not show the same statistically significant continence differences. Two less-experienced surgeons actually showed more efficient APM profiles than two more-experienced surgeons, demonstrating that surgical experience alone is not a reliable proxy for technical execution quality as measured by instrument kinematics.

This dissociation between experience and APM-measured efficiency is important: it suggests that some surgeons develop more efficient motion patterns earlier in their career, while others may accumulate experience without necessarily optimizing their technical execution -- a finding that has direct implications for surgical training and credentialing programs.

TL;DR: Surgeons clustered by APM efficiency showed statistically better early continence rates in a 493-case validation cohort, while traditional experience-based grouping failed to show the same differences -- confirming APMs measure a distinct dimension of surgical quality.
Pages 6-7
Implications for Surgeon Training and Patient Counseling

The finding that APMs outperform clinical features for predicting patient outcomes supports the development of objective, data-driven surgical credentialing systems that evaluate technique rather than relying solely on case volume thresholds or subjective mentor assessments.

Real-time or post-operative feedback systems built on APM analysis could guide surgeons toward more efficient motion patterns during the specific steps -- anastomosis and apical dissection -- that most strongly predict continence outcomes. This could accelerate skill development in training programs and support targeted remediation for surgeons whose patients show poor outcomes.

From a patient counseling perspective, a model that incorporates both surgical technique (APMs) and patient characteristics could provide more personalized estimates of continence recovery timelines, helping patients set realistic expectations and plan for rehabilitation or support services.

Key limitations of the study include the single-institution design, small surgeon sample (only 8 surgeons), and the relatively modest CI of 0.60, which reflects the inherent difficulty of predicting a complex biological recovery process. External validation across multiple institutions and surgeon cohorts will be required before clinical deployment.

TL;DR: APM-based outcome prediction lays the groundwork for objective surgical credentialing, real-time technique feedback, and personalized patient counseling, though external validation across multiple institutions is needed before clinical adoption.
Citation: Open Access, . Available at: PMC6706286.