review
Convolutional neural networks have served as the standard tool for computer vision tasks, yet they encounter difficulties when handling massive and complex datasets. To address these limitations, vision transformers adapt transformer models originally created for natural language processing to visual data. An overview of vision transformers outlines their foundational concepts, structural components, and architectural modifications. Different design approaches vary across metrics such as performance effectiveness and computational complexity. These systems support practical visual processing tasks, including image classification, object detection, and semantic segmentation across relevant real-world settings. Despite their potential to influence the wider field of computer vision, vision transformers present ongoing challenges and operational restrictions that define current research needs and future developmental directions.
Modern automated systems rely heavily on visual understanding to interpret images and video. By adapting models originally designed for language, vision transformers provide an alternative to traditional neural networks for analysing complicated visual information. Understanding how these models perform across different tasks helps guide the development of more capable and efficient visual recognition technologies.
Vision transformers apply to tasks such as object detection, image classification, and semantic segmentation, which are relevant to developers of computer vision software across varied industries. However, because this review analyses existing architectures, operational challenges, and ongoing research needs rather than deploying a specific tool, the technology represents early-stage to applied conceptual evaluation rather than an immediately packaged commercial product.
AI-generated from the published abstract. Always read the original work before citing.
In recent years, the development of deep learning has revolutionized the field of computer vision, especially the convolutional neural networks (CNNs), which become the preferred approach for numerous tasks handling images. However, CNNs have difficulty interpreting massive and complicated datasets, which has led to the creation of alternative architectures such as vision transformers. The transformer architecture, which was initially developed for natural language processing (NLP), is modified for image-related applications via vision transformers. In this paper, we present an outline of the main concepts and components of vision transformers. We review various variations and modifications to the architecture, and compare different approaches based on their effectiveness, complexity, and other attributes. Additionally, we examine the applications and uses of vision transformers, such as image classification, object detection, and semantic segmentation, and provide illustrations of relevant real-world situations. Finally, we discuss the potential impact of vision transformers on computer vision, while exploring the challenges and restrictions associated with their usage. We conclude by outlining potential new directions and advancements in the field of computer vision, as well as areas that require further study and investigation.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/cist56084.2023.10410015
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.