article · PLoS ONE
Accurate crop yield prediction is essential for agricultural planning, yet machine learning (ML) models remain highly sensitive to the quality and structure of input data. This study uses simulated datasets to systematically investigate how data structure (sample size and number of predictors), data imperfections (missing values and multicollinearity), and pre-processing methods (imputation techniques and principal component analysis) influence ML regression performance. The performance of ML-based models (RF, SVM, MLR, XGBoost, LightGBM, NNet, and kNN) on the pre-processed data was evaluated. In total, each algorithm was tested on 1,728 datasets and pre-processing scenarios, yielding 12,096 model-scenario evaluations across the seven ML algorithms. Results show that missing data decreases performance (R2 drops by up to 20.63%; MAE increases by 12.67%), while multicollinearity may inflate R2 values despite poorer MAE performance. Larger sample sizes consistently improve prediction accuracy (R2 = +18.87%; MAE = -8.16%), whereas more predictors generally reduce it (R2 = -8.47%; MAE = +44.74%). Regression-based imputation improved both R2 and MAE the most, while RF demonstrated greater robustness across varying conditions. These simulation findings were further validated on five real-world crop datasets (Maize, Yam, Cassava, Sorghum, and Peanuts), confirming that despite RF remains the safest and most robust default, no single pre-processing strategy is universally optimal and that the selection of imputation method and dimensionality reduction must be strictly contingent upon the dataset's specific missingness rate and correlation structure. Overall, this study highlights the complex interplay between data characteristics and pre-processing, urging the development of clearer guidelines for applying ML to agricultural datasets.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1371/journal.pone.0353938
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.