Assessment of PD-L1 Expression and Tumour Infiltrating Lymphocytes in Early-Stage Non-Small Cell Lung Carcinoma With Artificial Intelligence Algorithms

J Clin Pathol 2025 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
AI-Powered Pathology: Automating PD-L1 and TIL Scoring in Early NSCLC

The Biomarker Scoring Challenge PD-L1 immunohistochemistry (IHC) and tumor-infiltrating lymphocyte (TIL) assessment are critical biomarkers for immunotherapy selection in NSCLC. However, manual scoring is time-consuming, subject to significant inter-observer variability (Cohen's kappa 0.45-0.67), and dependent on pathologist experience. AI algorithms offer the potential for faster, more reproducible scoring.

Two AI Tools Evaluated Published in the Journal of Clinical Pathology in 2024, this prospective pilot study evaluated two commercially available AI pathology platforms - PathAI (AIM-PD-L1-NSCLC algorithm) and Navify Digital Pathology (NDP) by Roche - for PD-L1 and TIL scoring in 50 early-stage NSCLC specimens. Both tools were compared against manual scoring by experienced pathologists.

Study Significance Unlike AI development studies, this work evaluates commercial tools already available for potential clinical use. The findings directly inform whether pathology departments can currently trust these AI tools for clinical biomarker reporting, and whether AI-assisted scoring would change clinical treatment decisions compared to manual assessment.

TL;DR: This prospective study evaluated two commercial AI pathology tools (PathAI and Navify) against expert pathologist manual scoring for PD-L1 and TIL assessment in 50 early NSCLC cases, finding AI significantly faster with higher PD-L1 detection rates.
Pages 2-3
Study Cohort, Staining Protocol, and AI Algorithm Details

Patient Cohort and Tissue Preparation Fifty consecutive early-stage NSCLC specimens (40 adenocarcinomas, 9 squamous cell carcinomas, 1 adenosquamous carcinoma) were collected at Hospital Universitario 12 de Octubre, Madrid. All material was formalin-fixed paraffin-embedded (FFPE). PD-L1 IHC was performed with anti-PD-L1 clone SP263 on a VENTANA BenchMark ULTRA instrument, and CD8 IHC with clone SP57 for TIL assessment.

Manual PD-L1 Scoring Two pathologists with different experience levels (general pathologist with 10 years experience and thoracic pathologist with 20 years experience) independently scored PD-L1 tumor proportion score (TPS) as a continuous variable and in three categories: less than 1%, 1-49%, and 50% or greater. A third thoracic pathologist with 30 years experience provided consensus scores. Turn-around time (TAT) was recorded in seconds for all readers.

AI Algorithm Details The Navify Digital Pathology (NDP) SP263 algorithm classifies tumor cells as PD-L1 positive or negative using a supervised machine learning model, requiring the pathologist to draw a region of interest. The PathAI AIM-PD-L1-NSCLC algorithm uses convolutional neural networks and is semi-automated - the pathologist only excludes the external positive control. Both were assessed on the same scanned whole-slide images from a Roche Ventana DP200 scanner at 200x magnification.

TL;DR: Fifty FFPE NSCLC specimens were scored by two pathologists using manual PD-L1 SP263 IHC and by NDP and PathAI algorithms; CD8 IHC and H&E sections were assessed for TILs using NDP software and PathAI's tumor cellularity algorithm.
Pages 3-4
PD-L1 Scoring: AI Detects Higher Positivity Rates Than Manual Assessment

Key PD-L1 Finding Both AI tools identified significantly more cases at or above the 1% PD-L1 TPS threshold compared to manual scoring by either pathologist (p = 0.00015). Manual consensus classified 34% of cases as PD-L1 less than 1%, while both NDP and PathAI classified only approximately 10% as less than 1%. This suggests AI tools are more sensitive at detecting low-level PD-L1 positivity that human eyes may categorize as negative.

PD-L1 TPS Distribution Manual consensus median PD-L1 TPS was 2% (range 0-98%), while NDP consensus median was 4% (range 0-83%) and PathAI automated median was 3% (range 0-95%). The systematic shift upward in AI-reported TPS is most pronounced in the less than 1% vs. 1-49% threshold, which determines whether patients qualify for ICI monotherapy in many treatment guidelines.

Turn-Around Time Advantage TAT for AI-assisted PD-L1 scoring was significantly shorter than manual scoring (p less than 0.0001 for all comparisons). The PathAI algorithm was particularly fast as it required minimal pathologist interaction beyond excluding the external control. Exact TAT values were not reported in the abstract but the magnitude of difference was described as highly significant.

TL;DR: AI tools identified significantly more cases as PD-L1-positive (at or above 1% TPS) compared to manual scoring, with both NDP and PathAI classifying only 10% as PD-L1-negative versus 34% by manual consensus, while completing scoring significantly faster.
Pages 4-5
TIL Assessment and Correlation Between Methods

TIL Correlation Finding Total TIL density measured by the PathAI algorithm (from H&E slides) and total CD8+ cell density measured by NDP (from CD8 IHC slides) were significantly correlated (Kendall's tau = 0.49, 95% CI 0.37-0.61, p less than 0.0001). This moderate-to-strong correlation confirms that H&E-based TIL density can serve as a practical surrogate for CD8 IHC-based TIL quantification.

Practical Implications of H&E TIL Equivalence CD8 IHC requires an additional tissue section and staining procedure beyond routine H&E, adding cost and time. If PathAI's H&E-based TIL assessment correlates well with CD8 IHC counts, pathology departments could provide TIL data from the H&E slide already prepared for histological diagnosis, without ordering additional CD8 immunostaining.

Inter-Observer Agreement Patterns Agreement between observers was generally higher for tumors with low PD-L1 expression (less than 1%) than for intermediate TPS ranges (1-49%), regardless of whether the method was manual or AI. High-PD-L1 cases (50% or greater) also showed strong agreement. The intermediate range remains the most challenging for all methods, representing the key area where additional accuracy would have the greatest clinical impact.

TL;DR: H&E-based TIL density (PathAI) correlated significantly with CD8 IHC TIL counts (tau = 0.49), suggesting H&E TIL quantification could replace dedicated CD8 IHC in routine practice, while agreement was lowest in the PD-L1 1-49% range for all methods.
Pages 5-6
Implications of Higher AI-Detected PD-L1 Positivity for Clinical Decision-Making

Therapeutic Implications The 1% PD-L1 TPS threshold determines eligibility for pembrolizumab monotherapy in first-line NSCLC per current treatment guidelines. If AI tools systematically reclassify approximately 24% of cases from PD-L1-negative to PD-L1-positive (from 34% to 10% negative), a substantial proportion of patients who are currently denied ICI monotherapy might qualify under AI scoring. This reclassification has major therapeutic and economic implications.

Analytical vs. Clinical Sensitivity The higher AI-reported positivity could reflect superior detection of low-level PD-L1 staining that pathologists round down to negative, or it could reflect AI overcounting due to non-specific staining or artifact in challenging regions. Further correlation with clinical outcomes - whether AI-reclassified PD-L1-positive patients respond to ICIs - is needed to determine which interpretation is correct.

Standardization Across AI Platforms The study found that PathAI and NDP agreed with each other more than either agreed with manual scoring in certain subgroups, suggesting that AI tools may have a systematic analytical difference from manual pathology in how they interpret borderline staining intensity. This systematic difference needs to be characterized across tissue types and clone-platform combinations before AI scores can be interchangeably used with manual scores.

TL;DR: AI's higher PD-L1 positivity detection has direct therapeutic implications if accurate - reclassifying up to 24% of manual PD-L1-negative cases as positive - but requires clinical outcome correlation to confirm whether AI-detected positivity predicts ICI response.
Pages 6-7
Limitations and Future Research Priorities

Small Pilot Study With only 50 cases, this study is explicitly presented as a preliminary investigation. Statistical conclusions about AI versus manual agreement are limited by sample size, and the subgroup analyses by histological type (adenocarcinoma vs. squamous) lack statistical power. A larger validation study of several hundred cases is the immediate research priority.

Absence of Clinical Outcome Data No survival, response, or recurrence data were available for the 50 patients, preventing analysis of whether AI-scored PD-L1 or TIL values better predict treatment outcomes than manual scores. The most important validation study would link AI-measured biomarker values to actual ICI response in a prospective cohort.

Expansion to Multimarker AI Panels The study evaluated PD-L1 and TILs separately. Future AI pathology tools should provide integrated multi-biomarker profiling from a single scan - simultaneously quantifying PD-L1 TPS, TIL density, CD8 spatial distribution, and potentially TMB from NGS-annotated slides - to generate composite ICI eligibility scores that outperform individual markers.

TL;DR: This pilot study needs scale-up to hundreds of cases with linked clinical outcomes, and the next generation of AI pathology tools should provide integrated multi-biomarker profiling from single slides to move beyond single-marker assessment.
Citation: Open Access, 2025. Available at: PMC12322406.