article · Discover Public Health
Abstract The COVID-19 pandemic has posed significant challenges to developing countries like Nigeria due to limited resources. Accurate prediction of disease spread is crucial for effective containment measures. This study investigates the application of statistical and machine learning (ML) techniques in modelling and predicting COVID-19 cases in Nigeria, using data from January 2020 through December 2021. By analyzing demographic data (age, gender, location), symptom patterns, and contact tracing information, we seek to identify correlations and temporal trends associated with disease transmission. The datasets, obtained from the National Centre for Disease Control (NCDC), were cleaned before statistical analyses were carried out with Pearson’s Correlation, Analysis of Variance, and Cramer’s V Correlation. Prediction was carried out using the random forest (RF) classification model, implemented in Python's scikit learn library. Key findings include (1) 94.97% of confirmed contacts tested positive, underscoring high transmission rates; (2) occupations like healthcare workers and students were high-risk groups; and (3) the RF model achieved 87% accuracy in classifying source cases, though it struggled with minority classes. These can inform evidence-based policymaking and contribute to mitigating the impact of future outbreaks. A limitation of this study is the dependence on the accuracy of the NCDC data.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1186/s12982-026-02745-w
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.