Predicting Pancreatic Cancer in New-Onset Diabetes Cohort Using a Novel Model With Integrated Clinical and Genetic Indicators: A Large-Scale Prospective Cohort Study

Cancer Medicine 2024 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Page [1, 2]
Finding Pancreatic Cancer Early by Screening Diabetic Patients

People who develop new-onset diabetes (NOD) have nearly 8 times the risk of being diagnosed with pancreatic cancer compared to the general population. This makes new-onset diabetes a powerful early warning signal — but not all new diabetics have cancer; the vast majority have type 2 diabetes mellitus (T2DM) that is unrelated to a pancreatic tumor.

The challenge is accurately identifying which newly diabetic patients actually have underlying pancreatic cancer, versus those developing ordinary type 2 diabetes. This study used a large UK cohort and combined clinical data with genetic information to build a model capable of making that distinction.

TL;DR: New-onset diabetes is an 8-fold risk signal for pancreatic cancer, and this study built a machine learning model combining clinical and genetic data to identify which new diabetics actually have pancreatic cancer.
Pages 3-3
A Large-Scale Study Using UK Biobank and Genetic Markers

The study used the UK Biobank, a large prospective cohort of 502,407 participants followed over many years. From this group, 12,735 individuals developed new-onset diabetes during follow-up, and 100 of them were subsequently diagnosed with pancreatic cancer within 2 years of their diabetes diagnosis.

Researchers tested 82 clinical variables including age, blood tests, inflammatory markers, and lifestyle factors, alongside genetic data from 488,000 participants including over 24 single nucleotide polymorphisms (SNPs) identified from genome-wide association studies. Five different machine learning algorithms were compared for performance.

TL;DR: The study analyzed 12,735 new-onset diabetes cases from UK Biobank, using 82 clinical variables and 24 genetic markers with 5 machine learning algorithms to identify pancreatic cancer.
Pages 5-5
A Highly Accurate Model Combining Clinical and Genetic Information

The best-performing model was logistic regression integrating both clinical and genetic markers, achieving an AUC of 0.897 (95% CI: 0.865-0.929). The final clinical component used just 5 predictors: age at recruitment, platelet count, systolic blood pressure, immature reticulocyte fraction, and platelet crit. These combined with 24 genetic SNPs to create the full predictive model.

Among the clinical variables, age had the strongest individual predictive power. The 24 SNPs added substantially to clinical prediction alone, and the combined model significantly outperformed either clinical or genetic data used independently. This confirms that genetic and clinical information capture different and complementary aspects of cancer risk.

TL;DR: The combined clinical-genetic model achieved an AUC of 0.897, significantly outperforming clinical or genetic data alone in distinguishing pancreatic cancer from ordinary type 2 diabetes.
Page [6, 7]
Practical Cut-offs That Make Screening Feasible

The study established two practical risk threshold cut-offs. At the recommended 1.28% threshold, only 13% of new-onset diabetics would need further testing, but this approach captures 76% of all pancreatic cancer cases. This is a substantial efficiency improvement over current practice where all new-onset diabetics might be considered for screening.

A higher-risk threshold at 5.26% would require only 2% of the population to undergo definitive testing while still capturing nearly half of all cancers. Clinical teams could choose between these thresholds based on healthcare resources and patient preferences, giving the model practical flexibility for real-world implementation.

TL;DR: Two risk thresholds allow flexible clinical deployment: the recommended 1.28% threshold tests only 13% of new diabetics while capturing 76% of pancreatic cancer cases.
Pages 8-8
The Genetic Markers That Distinguish Cancer From Ordinary Diabetes

A particularly noteworthy finding was the identification of specific SNPs that differ between patients with pancreatic cancer-associated new-onset diabetes versus those with ordinary type 2 diabetes. This represents the first systematic attempt to use genetic sequencing to distinguish these two diagnoses in the new-onset diabetes setting.

This genomic layer of information is important because clinical symptoms and blood tests often cannot reliably distinguish the two conditions in the crucial early window when treatment of pancreatic cancer would be most effective. Genetic risk scores that are stable from birth could eventually be calculated for all adults, with high-risk individuals flagged automatically when diabetes develops.

TL;DR: The study identified specific genetic markers that distinguish pancreatic cancer-associated diabetes from ordinary type 2 diabetes, a world-first finding that could transform early detection.
Page [9, 10]
A Blueprint for Targeted Pancreatic Cancer Screening

This study demonstrates that integrating routine clinical data with genetic information in a machine learning framework can identify pancreatic cancer cases within the new-onset diabetes population with high accuracy. The UK Biobank cohort is predominantly white European, so validation in more ethnically diverse populations is needed before global deployment.

As genetic testing becomes cheaper and more routine, incorporating polygenic risk scores for pancreatic cancer into diabetes management could become standard practice. A patient newly diagnosed with diabetes who also has a high genetic risk score for pancreatic cancer could be immediately directed for targeted imaging surveillance, potentially catching the cancer at a curable stage.

TL;DR: Combining clinical and genetic data can reliably identify pancreatic cancer among new-onset diabetics, offering a blueprint for targeted early-detection screening programs in this high-risk population.
Citation: Open Access, 2024. Available at: PMC11551786.