review · Brain Informatics
Alzheimer's disease (AD), a significant public health challenge, requires accurate early diagnosis to improve patient outcomes. Vision Transformers (ViTs) and Convolutional Vision Transformers (CViTs) have emerged as powerful Deep Learning architectures for this task. Following PRISMA guidelines, this systematic review analyzes 68 studies selected from 564 publications (2021-2025) across five major databases: Scopus, Web of Science, ScienceDirect, IEEE Xplore, and PubMed. We introduce novel taxonomies to systematically categorize these works by model architecture, data modality, fusion strategy, and diagnostic objective. Our analysis reveals key trends, such as the rise of hybrid CViT frameworks, and critical gaps, including a limited focus on Mild Cognitive Impairment-to-AD progression. Critically, we also assess practical implementation details, revealing widespread challenges in algorithmic reproducibility. The discussion culminates in a forward-looking analysis of Large Vision Models and proposes future directions emphasizing the need for robust multimodal integration, lightweight transformer designs, and Explainable AI to advance AD research and bridge the critical gap between high-performance modeling and clinical applicability.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1186/s40708-025-00286-7
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.