People who develop new-onset diabetes (NOD) have nearly 8 times the risk of being diagnosed with pancreatic cancer compared to the general population. This makes new-onset diabetes a powerful early warning signal — but not all new diabetics have cancer; the vast majority have type 2 diabetes mellitus (T2DM) that is unrelated to a pancreatic tumor.
The challenge is accurately identifying which newly diabetic patients actually have underlying pancreatic cancer, versus those developing ordinary type 2 diabetes. This study used a large UK cohort and combined clinical data with genetic information to build a model capable of making that distinction.
The study used the UK Biobank, a large prospective cohort of 502,407 participants followed over many years. From this group, 12,735 individuals developed new-onset diabetes during follow-up, and 100 of them were subsequently diagnosed with pancreatic cancer within 2 years of their diabetes diagnosis.
Researchers tested 82 clinical variables including age, blood tests, inflammatory markers, and lifestyle factors, alongside genetic data from 488,000 participants including over 24 single nucleotide polymorphisms (SNPs) identified from genome-wide association studies. Five different machine learning algorithms were compared for performance.
The best-performing model was logistic regression integrating both clinical and genetic markers, achieving an AUC of 0.897 (95% CI: 0.865-0.929). The final clinical component used just 5 predictors: age at recruitment, platelet count, systolic blood pressure, immature reticulocyte fraction, and platelet crit. These combined with 24 genetic SNPs to create the full predictive model.
Among the clinical variables, age had the strongest individual predictive power. The 24 SNPs added substantially to clinical prediction alone, and the combined model significantly outperformed either clinical or genetic data used independently. This confirms that genetic and clinical information capture different and complementary aspects of cancer risk.
The study established two practical risk threshold cut-offs. At the recommended 1.28% threshold, only 13% of new-onset diabetics would need further testing, but this approach captures 76% of all pancreatic cancer cases. This is a substantial efficiency improvement over current practice where all new-onset diabetics might be considered for screening.
A higher-risk threshold at 5.26% would require only 2% of the population to undergo definitive testing while still capturing nearly half of all cancers. Clinical teams could choose between these thresholds based on healthcare resources and patient preferences, giving the model practical flexibility for real-world implementation.
A particularly noteworthy finding was the identification of specific SNPs that differ between patients with pancreatic cancer-associated new-onset diabetes versus those with ordinary type 2 diabetes. This represents the first systematic attempt to use genetic sequencing to distinguish these two diagnoses in the new-onset diabetes setting.
This genomic layer of information is important because clinical symptoms and blood tests often cannot reliably distinguish the two conditions in the crucial early window when treatment of pancreatic cancer would be most effective. Genetic risk scores that are stable from birth could eventually be calculated for all adults, with high-risk individuals flagged automatically when diabetes develops.
This study demonstrates that integrating routine clinical data with genetic information in a machine learning framework can identify pancreatic cancer cases within the new-onset diabetes population with high accuracy. The UK Biobank cohort is predominantly white European, so validation in more ethnically diverse populations is needed before global deployment.
As genetic testing becomes cheaper and more routine, incorporating polygenic risk scores for pancreatic cancer into diabetes management could become standard practice. A patient newly diagnosed with diabetes who also has a high genetic risk score for pancreatic cancer could be immediately directed for targeted imaging surveillance, potentially catching the cancer at a curable stage.