Pancreatic ductal adenocarcinoma (PDAC) is typically diagnosed at an advanced stage, when treatment options are limited. A major barrier to earlier diagnosis is the absence of a reliable, non-invasive screening test. Blood tests like CA19-9 are widely used but are insufficiently sensitive to catch the cancer early, and they miss a significant fraction of tumors.
Urine offers an attractive alternative to blood. Urine is easy to collect, requires no needles, and contains proteins and other molecules shed by the kidneys, bladder, and—importantly—tumor cells throughout the body. Several studies have identified specific proteins in urine that are elevated in pancreatic cancer patients.
This study assembled a panel of urinary biomarkers previously linked to PDAC and tested whether machine learning could combine them into a more accurate diagnostic tool than any single biomarker alone. The researchers then used single-cell transcriptomic data as an independent biological validation.
The research team assembled a dataset of urinary biomarker measurements from PDAC patients, patients with pancreatitis (a benign pancreatic condition that can mimic cancer symptoms), and healthy controls. Key biomarkers included REG1A, REG1B, TFF1, LYVE1, and plasma CA19-9 (measured alongside the urinary panel).
Two distinct classification tasks were designed. Binary classification distinguished PDAC from all non-PDAC cases—simulating a yes/no cancer screening scenario. Multiclass classification separated PDAC, pancreatitis, and healthy control into three groups—a more clinically realistic challenge, since distinguishing cancer from pancreatitis is notoriously difficult.
Multiple machine learning algorithms were compared, including logistic regression, support vector machines, random forests, XGBoost, and deep learning architectures. Missing values in biomarkers (a practical reality in clinical data) were handled using multiple imputation techniques, and cross-validation was used to estimate model performance without overfitting.
For the simpler binary task (PDAC vs. non-PDAC), deep learning achieved the highest accuracy at 91%. For the more challenging three-way classification (PDAC vs. pancreatitis vs. healthy), XGBoost performed best with 87% accuracy—remarkably high given the difficulty of distinguishing cancer from pancreatitis, which produces many overlapping symptoms and biomarker elevations.
REG1A emerged as one of the most important biomarkers in the panel—a protein found in the digestive system that is dramatically elevated in the urine of PDAC patients but not in healthy individuals or pancreatitis patients. LYVE1 and TFF1 also contributed meaningfully to model predictions.
The combination of biomarkers consistently outperformed any single biomarker used alone, confirming the premise that a multi-marker AI-driven approach captures disease biology more completely than CA19-9 or any other individual measurement.
To validate that the biomarkers reflect real tumor biology rather than statistical coincidence, the researchers performed a 'post hoc' validation using single-cell transcriptomic data. This technique measures gene activity in individual cells, allowing researchers to identify which cell types within pancreatic tumors actually produce the proteins detected in urine.
The analysis confirmed that genes encoding REG1A, LYVE1, TFF1, and other key biomarkers are specifically upregulated in pancreatic cancer cells and associated stromal cells—not in normal pancreatic tissue or pancreatitis cells. This biological grounding strengthens confidence that the urine test is detecting genuine cancer-related signals.
Single-cell data also helped identify which cell populations shed the most biomarker protein into the urine, providing insight into the cancer biology driving the test's performance. Ductal cancer cells and cancer-associated fibroblasts appeared to be the primary sources.
A urine test for pancreatic cancer has obvious practical appeal: it requires no blood draws, is comfortable for patients, can be performed at home or in any clinical setting, and can be repeated frequently for surveillance. If the biomarker panel validated here can be commercialized as a standardized assay, it could serve as a first-line screen for high-risk groups.
The panel's ability to differentiate PDAC from pancreatitis is particularly clinically valuable. Pancreatitis is common in patients presenting with abdominal pain, and distinguishing it from early PDAC currently requires expensive imaging or endoscopic procedures. A urine test achieving 87% accuracy for this distinction could substantially reduce unnecessary invasive workups.
The next steps include validating the test in independent prospective cohorts, optimizing the biomarker panel composition, and developing a standardized measurement assay. Clinical utility trials—testing whether use of the test actually leads to earlier diagnoses and improved outcomes—will be the ultimate measure of success.
This study demonstrates that a machine learning framework applied to urinary biomarkers can achieve clinically meaningful accuracy for both detecting pancreatic cancer and distinguishing it from pancreatitis. The 91% binary accuracy and 87% three-way accuracy represent substantial advances over single-biomarker approaches.
The integration of single-cell transcriptomic validation is a methodological strength that distinguishes this study from purely statistical biomarker analyses. By anchoring the biomarkers in tumor cell biology, the study strengthens the scientific case for their clinical use.
If translated into a clinical product, this approach could shift pancreatic cancer detection toward a regular, painless screening paradigm—potentially catching tumors years earlier than current methods, when surgery remains possible and cure is achievable.