A Systematic Review and Meta-Analysis of Artificial Intelligence Versus Clinicians for Skin Cancer Diagnosis

NPJ Digit Med 2024 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
How Well Does AI Diagnose Skin Cancer Compared to Doctors?

The AI Revolution in Dermatology Artificial intelligence research in dermatology has grown exponentially, with dozens of studies claiming AI matches or exceeds dermatologist performance. Yet a rigorous quantitative synthesis was lacking. This systematic review and meta-analysis evaluated 53 comparative studies to provide an evidence-based answer to how AI actually performs against clinicians.

Study Design Following PRISMA guidelines, researchers searched PubMed, Embase, and Cochrane Library through August 2022. Studies had to compare AI algorithm performance directly against clinicians using dermoscopic or clinical images. The quality of each study was assessed using the validated QUADAS-2 tool, and 19 studies met the stricter criteria for formal meta-analysis.

Why This Matters Several AI systems for skin cancer detection have already received CE approval in Europe and are entering clinical practice. Understanding the actual evidence base - including limitations and gaps - is critical for responsible deployment and for guiding future research priorities.

TL;DR: This meta-analysis pooled data from 53 studies to rigorously compare AI algorithm performance against dermatologists and other clinicians for skin cancer diagnosis.
Pages 2-3
Meta-Analysis Design and Quality Assessment

Inclusion Criteria Studies were included if they directly compared AI performance to clinicians, used skin lesion images as the input, and reported extractable diagnostic accuracy data. For the meta-analysis, additional requirements included dermoscopic images only, binary classification (benign/malignant or melanoma/nevus), and clear stratification of clinician expertise level.

Clinician Expertise Categories Clinicians were divided into expert dermatologists (fellowship-trained specialists), non-expert dermatologists (residents or general dermatologists without dermoscopy specialization), and generalists (non-dermatology physicians). This stratification allowed the meta-analysis to test whether AI performance advantage differed depending on who it was compared to.

Risk of Bias Assessment Most studies (58%) had uncertain risk of bias under QUADAS-2 evaluation. Only 26% had low risk. Common methodological weaknesses included use of curated image sets not representative of real clinical populations, and lack of external validation. The meta-analysis separately analyzed results from internal versus external test sets, a distinction with major implications for generalizability.

TL;DR: Nineteen of 53 studies met rigorous inclusion criteria for meta-analysis; most studies used dermoscopic images and were rated as having uncertain quality by independent reviewers.
Pages 3-5
AI Outperforms Clinicians Overall, but the Gap Depends on Expertise

Overall Performance Comparison Pooled across all 19 meta-analysis studies, AI algorithms achieved sensitivity of 87.0% and specificity of 77.1%. All clinicians combined achieved sensitivity of 79.8% and specificity of 73.6%. Both differences were statistically significant, indicating AI has a genuine - if modest - overall advantage in controlled study settings.

AI vs. Expert Dermatologists When compared only to expert dermatologists, the gap narrowed substantially. AI achieved sensitivity 86.3% and specificity 78.4%, versus expert dermatologists at sensitivity 84.2% and specificity 74.4%. These differences were described as 'clinically comparable,' meaning expert dermatologists perform very similarly to AI in head-to-head comparisons.

AI vs. Generalists The largest performance gap emerged between AI and generalist physicians. AI had sensitivity 92.5% versus generalists' 64.6%, and specificity 66.5% versus 72.8%. AI shows markedly better sensitivity but lower specificity than generalists, suggesting AI could be most valuable as a screening tool for non-specialist settings.

TL;DR: AI outperforms clinicians overall, but nearly matches expert dermatologists. The largest advantage is over generalist physicians, where AI sensitivity was 28 percentage points higher.
Pages 4-5
Internal Test Sets Inflate AI Performance - External Validation Tells a Different Story

The Validation Problem A critical finding was that AI performance was substantially higher in studies using internal test sets (data from the same source as training data) compared to external test sets (data from different institutions or countries). This pattern suggests significant overfitting and lack of generalizability in many published studies.

Clinical Significance For a diagnostic AI to be clinically useful, it must perform well on external data - patients from different hospitals, demographics, and imaging equipment. The gap between internal and external performance in these studies is a warning sign that many published AI claims may not translate to real-world practice.

Implications for AI Development Future studies should prioritize external validation on diverse, real-world patient populations. Prospective studies - which were underrepresented in this review - are needed to understand how AI performs when integrated into actual clinical workflows rather than controlled research settings.

TL;DR: AI performance in studies using internal test sets was significantly higher than in external validation studies, raising concerns that published results may overstate real-world clinical utility.
Pages 5-6
AI Assistance Improves Clinician Performance

Human-AI Collaboration Several included studies examined how clinician performance changed when given AI support rather than simply comparing them head-to-head. When dermatologists used AI assistance, their diagnostic accuracy generally improved - particularly their sensitivity for detecting melanoma. This suggests the most valuable use of AI may be as a collaborative tool rather than a replacement.

Augmentation vs. Replacement The meta-analysis authors note that the framing of 'AI vs. clinicians' may be less relevant than 'AI plus clinicians.' Real-world deployment should focus on how AI tools integrate into clinical workflows, and which specific decision points benefit most from algorithmic support.

Confidence and Low-Confidence Cases Some studies found that AI was most helpful for cases where clinicians reported low confidence. Substituting AI classifications for clinician evaluations specifically in uncertain cases improved overall accuracy - pointing toward a practical implementation strategy where AI serves as a safety net for difficult cases.

TL;DR: When used as a support tool rather than a replacement, AI improved clinician performance, especially for difficult cases where human confidence was low.
Pages 6-7
Limitations of Current Evidence

Non-Representative Datasets Most studies used curated, high-quality image datasets that do not reflect the full spectrum of lesions encountered in clinical practice. Images with poor quality, ambiguous diagnoses, or rare presentations are often excluded from AI training sets, potentially inflating performance metrics.

Lack of Clinical Context Studies typically provided only the dermoscopic image without clinical metadata such as patient age, skin phototype, lesion history, or anatomic location. In real clinical practice, this context significantly influences diagnostic decisions and may give AI a disadvantage in studies that include it.

Heterogeneity and Publication Bias Significant heterogeneity existed across studies in terms of datasets, diagnostic tasks, and study designs. There was likely publication bias toward positive results. The majority of AI studies focused on melanoma versus nevus - a relatively narrow diagnostic task - limiting conclusions about AI for the full range of skin cancers.

TL;DR: Most published AI dermatology studies use curated datasets and lack clinical context, limiting how well their findings translate to actual clinical practice.
Pages 7-8
Where AI in Dermatology Needs to Go Next

Real-World Prospective Studies The field urgently needs prospective studies that evaluate AI performance in real clinical settings, including consecutive unselected patients. Such studies should measure not just accuracy metrics but patient outcomes, workflow efficiency, clinician burden, and healthcare economics.

Diverse and Representative Datasets AI algorithms must be trained and tested on datasets that include diverse skin phototypes, imaging equipment, clinical settings, and demographic groups. Current training datasets are heavily skewed toward Fitzpatrick skin types I-III, raising equity concerns about performance in patients with darker skin tones.

Regulatory and Ethical Frameworks As AI systems receive regulatory approval and enter clinical use, robust post-market surveillance is needed. Transparent reporting standards, explainability requirements, and clear protocols for human oversight of AI-flagged cases will be essential for safe integration into dermatologic care.

TL;DR: Future AI dermatology research must shift from curated benchmark studies to prospective real-world trials using diverse populations, with clear frameworks for responsible clinical deployment.
Citation: Open Access, 2024. Available at: PMC11094047.