MARATTO

article · Diagnostics

Diagnostic Accuracy of Multimodal Large Language Models for Four-Class Benchmark of Oral Autoimmune Blistering Diseases: A Multicenter Paired Study

2026Open accessAin Shams University

Abstract

Background/Objectives: To compare the diagnostic performance of Claude Opus 4.7 and Gemini Pro 3 for the differential diagnosis of oral autoimmune blistering diseases (AIBDs) and evaluate their diagnostic reasoning, confidence, and calibration. Materials and Methods: This retrospective multicenter paired diagnostic accuracy study included 200 clinicopathologically confirmed AIBD cases (50 each of pemphigus vulgaris, mucous membrane pemphigoid, bullous pemphigoid, and linear IgA bullous dermatosis). Each case was independently assessed by both models using identical standardized clinical information and clinical photographic inputs. The task required forced-choice classification among the four predefined diseases. Histopathological and direct immunofluorescence findings were used exclusively to establish the clinicopathological reference diagnosis and were not provided to the AI models. The reference diagnosis was established by clinicopathological correlation. The primary outcome was diagnostic accuracy. Secondary outcomes included disease-specific diagnostic performance, Cohen’s κ, ROC analysis, calibration, confidence, reasoning quality, management recommendations, and error patterns. Pre-consensus inter-rater reliability of the two human assessors was also evaluated using Cohen’s κ for binary outcomes and weighted Cohen’s κ for the ordinal reasoning-quality score. Results: Claude achieved significantly higher diagnostic accuracy than Gemini (92.0% vs. 86.0%, p = 0.012), stronger agreement with the reference standard (κ = 0.893 vs. 0.813), and superior discrimination (macro-AUC 0.998 vs. 0.965). Claude demonstrated higher key diagnostic-feature identification (92.0% vs. 86.0%; p = 0.012) and higher clinical-reasoning scores (61.0% vs. 40.0% of responses rated good; Wilcoxon p < 0.001; r = 0.47), whereas management recommendations did not differ significantly (100.0% vs. 98.0%; p = 0.125). Calibration results were metric-dependent: Claude had a lower one-vs-rest Brier score (0.0468 vs. 0.0558), whereas Gemini had a lower expected calibration error (0.083 vs. 0.251). For both models, the predominant error was misclassification of linear IgA bullous dermatosis as mucous membrane pemphigoid. Conclusions: Both multimodal LLMs showed high performance in this controlled four-class benchmark, with Claude Opus 4.7 outperforming Gemini Pro 3 in overall accuracy and reasoning quality. However, these findings do not establish autonomous diagnostic capability, clinical effectiveness, or safety. The LABD–MMP misclassification and metric-dependent calibration highlight important limitations. The models should therefore be regarded as investigational adjunctive decision-support tools requiring clinician oversight and diagnostic verification. Prospective external and human-in-the-loop validation is required before clinical implementation.

Research topics

  • Autoimmune Bullous Skin Diseases
  • Dermatological and COVID-19 studies
  • Autoimmune and Inflammatory Disorders

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.3390/diagnostics16172761

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.