article
Diabetes is a disease that has no permanent cure, it is one of the most diseases leading to death in the world; hence early detection and diagnosis are required to prevent diabetes or lead to improve treatment. Machine learning has become increasingly used in the field of diseases. Our study aims to compare the performance of five different machine learning (ML) algorithms using accuracy as the primary metric. The study well be based on the Pima Indian Diabetes Database [1] to define the most effective algorithm for predicting diabetes. We hypothesize that the step of cleaning dataset by treating missing values can scale up the performance of models in diabetes prediction and diagnosis. In this study the dataset used consists of 768 instances, with 268 patients belong to the diabetic class and 500 patients belong to the non -diabetic class. The result obtained show that Random Forest (RF) model is more efficient at predicting diabetes compared to other applied algorithms achieving an accuracy of 80.5% using mean imputation method.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1145/3607720.3607764
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.