article · Journal of Smart Cities and Society
Data cleaning is a critical phase in data processing that detects and removes errors, duplicates, anomalies, and missing values to improve machine learning performance. Because model effectiveness relies heavily on data quality, substantial effort is dedicated to cleaning datasets before training. An examination of data cleaning facets and an evaluation of prominent outlier detection algorithms through benchmarking establishes the need for efficient, consistent detection methods. To enhance these outcomes, a new algorithm combines Isolation Forest with clustering techniques. By uniting the strengths of both approaches, the hybrid method seeks to improve outlier detection performance and refine broader data processing and analysis paradigms within contemporary data-driven environments.
Machine learning systems depend entirely on the quality of their training data. Flawed, duplicate, or anomalous records can distort automated predictions and decisions. Improving outlier detection through hybrid algorithmic methods helps organisations clean datasets more reliably, ensuring that downstream automated tools operate on sound and accurate information.
The algorithm could enhance automated data preparation software used by data scientists and software developers prior to model training. Because the abstract describes the proposal and benchmarking of an algorithm combining Isolation Forest with clustering without mentioning live deployment or product integration, the work appears to be early-stage research requiring further development into software libraries or tools.
AI-generated from the published abstract. Always read the original work before citing.
Data cleaning, also referred to as data cleansing, constitutes a pivotal phase in data processing subsequent to data collection. Its primary objective is to identify and eliminate incomplete data, duplicates, outdated information, anomalies, missing values, and errors. The influence of data quality on the effectiveness of machine learning (ML) models is widely acknowledged, prompting data scientists to dedicate substantial effort to data cleaning prior to model training. This study accentuates critical facets of data cleaning and the utilization of outlier detection algorithms. Additionally, our investigation encompasses the evaluation of prominent outlier detection algorithms through benchmarking, seeking to identify an efficient algorithm boasting consistent performance. As the culmination of our research, we introduce an innovative algorithm centered on the fusion of Isolation Forest and clustering techniques. By leveraging the strengths of both methods, this proposed algorithm aims to enhance outlier detection outcomes. This work endeavors to elucidate the multifaceted importance of data cleaning, underscored by its symbiotic relationship with ML models. Furthermore, our exploration of outlier detection methodologies aligns with the broader objective of refining data processing and analysis paradigms. Through the convergence of theoretical insights, algorithmic exploration, and innovative proposals, this study contributes to the advancement of data cleaning and outlier detection techniques in the realm of contemporary data-driven environments.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.3233/scs-230008
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.