Performance of a deep neural network in teledermatology: a single-centre prospective diagnostic study

J Eur Acad Dermatol Venereol 2021 AI 5 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Testing an AI Dermatology Tool in Real-World Teledermatology

Study context: This is one of the few prospective studies testing a commercial AI dermatology algorithm in an actual clinical teledermatology workflow, rather than in a controlled laboratory setting with curated images.

The AI system: A 174-class deep neural network (developed by a commercial provider) was evaluated on 340 cases submitted to a German university hospital's teledermatology service. Each case included one to several photographs and brief clinical notes.

Main findings: The AI's top-1 accuracy (did the correct diagnosis appear as the first suggestion?) was 41.2%, compared to 60.1% for attending dermatologists. For in-distribution cases where the AI had been trained on similar conditions, balanced accuracy improved to 47.6%.

Significance: While the AI's overall accuracy lagged behind dermatologists, its performance was comparable on in-distribution cases, and the study identified key factors - image quality and skin type - that substantially degraded AI performance in real-world use.

TL;DR: In a real teledermatology workflow, a 174-class AI algorithm reached 41.2% top-1 accuracy versus 60.1% for dermatologists, but matched specialist performance on cases within its training distribution.
Pages 2-3
How the Prospective Trial Was Structured

Patient population: 340 cases from 281 patients were submitted to the teledermatology service between April and October 2019. Cases represented the genuine referral mix, not a curated research set, including conditions ranging from inflammatory dermatoses to skin tumors.

AI evaluation protocol: Images from each teledermatology case were submitted to the AI system, which returned a ranked list of up to 10 diagnostic suggestions. Top-1 accuracy (correct diagnosis as first suggestion) and top-5 accuracy were measured.

Dermatologist comparison: The same cases were evaluated by senior dermatologists at the institution. Two separate reads were conducted: one with full clinical information and one with only images, to measure the contribution of context to diagnostic accuracy.

In-distribution vs. out-of-distribution: The 174-class AI was trained on certain skin condition categories. Cases involving conditions not well represented in training (out-of-distribution) were identified and analyzed separately to understand the AI's scope of competence.

TL;DR: The trial used 340 real teledermatology referrals to compare a 174-class AI to senior dermatologists, separately analyzing cases within and outside the AI's training distribution.
Pages 4-6
Accuracy, Limitations, and What Hurt Performance

Overall accuracy gap: The AI achieved top-1 accuracy of 41.2% vs. dermatologists' 60.1%. Top-3 accuracy for the AI was 63.7%, meaning the correct diagnosis appeared within the first three suggestions in nearly two-thirds of cases.

In-distribution performance: For the 221 cases within the AI's training distribution, balanced accuracy reached 47.6% - much closer to dermatologist-level performance and statistically comparable in some subgroup analyses.

Image quality impact: Cases with poor image quality (blurry, incorrectly framed, or dark photos typical of real-world smartphone submissions) showed substantially worse AI accuracy. Dermatologists were better able to compensate for poor images using clinical context.

Skin type effects: The AI performed less accurately on darker Fitzpatrick skin types (IV-VI). This reflects the well-documented underrepresentation of non-white skin in most dermoscopy and clinical photography training datasets.

TL;DR: The AI reached 41.2% top-1 accuracy overall but 47.6% balanced accuracy on in-distribution cases; image quality and skin type were the main factors limiting performance.
Pages 6-7
What This Means for AI in Teledermatology Practice

Triage vs. diagnosis: Even at 41-47% top-1 accuracy, the AI's top-3 list captured the correct diagnosis in 64% of cases. As a triage or prioritization tool that presents a differential for clinician review, this level of performance may still add value.

Workflow integration: The AI could flag high-risk conditions (melanoma, squamous cell carcinoma) for expedited review, even if it doesn't produce a definitive diagnosis. This could reduce time-to-specialist for urgent cases in teledermatology systems.

Human-AI synergy: Studies consistently show that AI-assisted dermatologists outperform unassisted dermatologists. The model should be positioned as a decision-support tool that prompts clinicians to consider diagnoses they might overlook, not as a standalone diagnostic.

Equity concerns: The performance gap on darker skin types highlights a fairness issue. Deploying AI tools that underperform in diverse populations could worsen existing health disparities in dermatology care.

TL;DR: Despite lower accuracy than dermatologists, the AI's top-3 suggestion list was correct in 64% of cases, supporting its use as a triage or differential-generation tool rather than a standalone diagnostic.
Pages 8-9
Improving Real-World AI Teledermatology Performance

Image quality standardization: Future teledermatology platforms should include real-time image quality feedback to guide non-specialist photographers in capturing adequate images before submission, reducing one of the biggest performance gaps.

Diverse training data: The performance drop on darker skin types will only be resolved by collecting and including large, well-annotated datasets from diverse global populations in model training and validation.

Out-of-distribution detection: AI systems should flag cases that fall outside their training distribution rather than producing low-confidence guesses. Uncertainty quantification methods could enable the AI to say 'I am not confident' and escalate for human review.

Multi-modal input: Incorporating structured clinical metadata (patient age, lesion duration, symptom history) alongside images could close part of the accuracy gap between AI and clinicians who naturally integrate this information in their diagnoses.

TL;DR: Better image quality standards, more diverse training data, uncertainty detection, and clinical metadata integration are the key steps to bring real-world AI teledermatology closer to specialist performance.
Citation: Open Access, 2021. Available at: PMC8274350.