article · Array
Modern machine learning models require substantial amounts of high-quality, annotated data for optimal performance. However, collecting and annotating this data is often a manual, time-consuming, and resource-intensive process, making it challenging to obtain sufficient training data for many real-world applications. Data augmentation is recognised as the most effective strategy to address this issue by increasing the volume, quality, and diversity of training data. This paper provides a comprehensive review of modern data augmentation methods specifically applicable to computer vision. It covers advanced techniques, including deeply learned, feature-level, and meta-learning-based strategies, as well as data synthesis approaches using 3D graphics, neural rendering, and generative adversarial networks. The survey also compares the performance of various state-of-the-art methods across different datasets and tasks.
Ensuring machine learning models perform well often requires vast amounts of data, which is costly and difficult to acquire. Data augmentation offers a critical solution by artificially expanding and diversifying datasets. This enables the development of more robust and accurate AI systems, particularly in fields like computer vision, without the prohibitive costs of manual data collection.
This research provides a foundational overview of techniques that are crucial for developing and improving machine learning applications, especially in computer vision. It directly supports engineers and researchers working on AI systems by guiding the selection and implementation of data augmentation strategies. While not a direct product, it informs the development of more effective AI solutions across various industries, making it relevant for organisations building or utilising AI-powered products and services. This is applied research, informing the practical application of existing and emerging techniques.
AI-generated from the published abstract. Always read the original work before citing.
To ensure good performance, modern machine learning models typically require large amounts of quality annotated data. Meanwhile, the data collection and annotation processes are usually performed manually, and consume a lot of time and resources. The quality and representativeness of curated data for a given task is usually dictated by the natural availability of clean data in the particular domain as well as the level of expertise of developers involved. In many real-world application settings it is often not feasible to obtain sufficient training data. Currently, data augmentation is the most effective way of alleviating this problem. The main goal of data augmentation is to increase the volume, quality and diversity of training data. This paper presents an extensive and thorough review of data augmentation methods applicable in computer vision domains. The focus is on more recent and advanced data augmentation techniques. The surveyed methods include deeply learned augmentation strategies as well as feature-level and meta-learning-based data augmentation techniques. Data synthesis approaches based on realistic 3D graphics modeling, neural rendering, and generative adversarial networks are also covered. Different from previous surveys, we cover a more extensive array of modern techniques and applications. We also compare the performance of several state-of-the-art augmentation methods and present a rigorous discussion of the effectiveness of various techniques in different scenarios of use based on performance results on different datasets and tasks.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.array.2022.100258
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.