MARATTO

dataset · Zenodo (CERN European Organization for Nuclear Research)

Crop Recommendation Dataset with SMOTE, VAE and CropGAN Synthetic Generated Dataset Variants

Abstract

Crop Recommendation Dataset with Synthetic Data Augmentation 1. Overview This dataset was developed to support research in data-driven crop recommendation systems. It contains soil and environmental attributes used to predict suitable crops for cultivation. In addition to the original dataset, this repository includes preprocessed data and synthetically generated samples produced using advanced data augmentation techniques. The dataset is designed to address class imbalance and improve model performance in precision agriculture applications. 2. Dataset Contents The repository includes the following files: Original_dataset.csvOriginal dataset collected from field survey Original_preprocessed_dataset.csvCleaned and transformed dataset used for model training (includes encoding and feature engineering) Smote_dataset.csvSynthetic dataset generated using SMOTE to address class imbalance Vae_dataset.csvSynthetic dataset generated using Variational Autoencoder (VAE) Cropgan_dataset.csvSynthetic dataset generated using the proposed GAN-based framework (CropGAN) Data_dictionary.csvDescription of all features and their meanings 3. Features Description Typical attributes include: Soil properties (e.g., pH, soil_texture, nutrients (N, P, K)) Environmental factors (e.g., rainfall, temperature, humidity) Crop label (target variable) Categorical variables such as soil texture were encoded during preprocessing. 4. Data Collection The dataset used in this study was originally collected by Kwaghtyo et al. [6]. The initial dataset consisted of 2,200 samples. To enhance model robustness and enable comprehensive experimentation, the dataset was expanded to 5,000 samples. The original 2,200 samples were retained to ensure comparability with previous studies, while an additional 2,800 samples were collected using a random sampling approach across farm fields in Yandev District, Gboko Local Government Area (LGA), Benue State, Nigeria. Data collection for the additional 2,800 samples was conducted across two farming seasons (2024 and 2025) to capture seasonal variability. Soil samples were collected at a depth of 0-30 cm using standard soil auger techniques and analysed in the laboratory to determine soil nutrient composition and physicochemical properties. Climatic variables were also recorded, with temperature ranging from 25.0 °C to 33.5 °C and rainfall ranging from 900 mm to 1,200 mm. The dataset comprises nine attributes: nitrogen (N), phosphorus (P), potassium (K), soil pH, humidity, temperature, rainfall, soil texture, and crop label. The target crops include maize, rice, pepper, soybean, beans, orange, guinea corn, cassava, tomatoes, and yam, which are commonly cultivated in the study area. 5. Preprocessing Steps The following preprocessing steps were applied: Handling missing values Normalisation/standardisation One-hot encoding of categorical variables Feature selection (where applicable) 6. Synthetic Data Generation To mitigate class imbalance, the following techniques were applied: SMOTE (Synthetic Minority Oversampling Technique) Variational Autoencoder (VAE) Proposed GAN-based model (CropGAN) Each method generates additional samples while preserving the statistical properties of the original dataset. 7. Intended Use This dataset is intended for: Crop recommendation research Machine learning model development Class imbalance studies Synthetic data evaluation 8. Limitations The dataset may not generalise to all geographic regions Synthetic data may introduce minor distributional deviations Results depend on preprocessing and model configuration 9. Reproducibility All datasets used in the study (original, original_preprocessed, and the synthetic variants) are provided to ensure reproducibility of experimental results. 10. Citation If you use this dataset, please cite: Kwaghtyo, K.D., Eke, C.I. and Ajon, A.T. (2026). Crop Recommendation Dataset with Synthetic Data Augmentation. 11. Contact For questions or collaboration:Dekera Kenneth KwaghtyoFederal University of Lafia, PMB 146, Nasarawa State, Nigeriadekera.kwaghtyo@student.fulafia.edu.ng

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.19709807

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.