article · Results in Engineering
• We present the first comprehensive evaluation of BEDL UQ method in medical domain. • BEDL is benchmarked across three cancer types and imaging modalities under a multidimensional evaluation setting . • BEDL outperforms MCD and VI Bayesian methods in classification performance, calibration, and robustness to distribution shifts. • BEDL achieves competitive OOD detection and supports uncertainty-guided rejection for safe clinical decision-making. • BEDL offers a computationally efficient and scalable alternative to conventional Bayesian methods. Deep learning models have shown promising capabilities in medical image analysis, yet their safe and trustworthy deployment in clinical oncology remains limited. One major barrier is the lack of reliable uncertainty quantification (UQ), as most models provide deterministic predictions, often exhibit overconfidence, and fail to recognize when they are uncertain—essentially, they don’t know what they don’t know. To address this issue, we investigate the uncertainty quantification capabilities of Bayesian Evidential Deep Learning (BEDL) for medical image classification. We present a comprehensive evaluation of this approach for cancer classification across three clinically relevant imaging modalities: histology for ovarian cancer, cytology for cervical cancer, and ultrasound for breast cancer. BEDL is benchmarked against two widely adopted Bayesian UQ methods—Monte Carlo Dropout (MCD) and Variational Inference (VI)—in diverse evaluation settings, including in-distribution testing, distribution shift, and out-of-distribution (OOD) detection. We systematically analyze each method's classification performance, calibration quality, robustness to distribution shifts, and the reliability of their uncertainty estimates. Results showed that BEDL achieves competitive classification performance across datasets and often provides improved calibration compared to MCD, with performance comparable to VI in many settings. Under distribution shift and OOD detection scenarios, the relative performance of the methods varies across datasets and uncertainty measures, highlighting the importance of dataset-specific evaluation. In addition, BEDL demonstrated favourable computational efficiency compared to VI and enables effective uncertainty-guided rejection for selective prediction. Overall, our findings suggest that BEDL represents a promising alternative for uncertainty-aware medical image analysis and highlight the importance of carefully evaluating uncertainty modelling approaches across diverse clinical scenarios to support the development of trustworthy AI systems in oncology.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.rineng.2026.110453
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.