Data Modeling

Mathematical and Computational Data Modeling
1
Citations
7k
Views
37
Articles
Your new experience awaits. Try the new design now and help us make it even better
Switch to the new experience
RESEARCH ARTICLE   (Open Access)

Predicting Prediabetes from Simple Clinical Questionnaires: A Machine Learning Approach Using K-Nearest Neighbors and Decision Tree Models

Abstract 1. Introduction 2. Methods 3. Results 4. Discussion 5. Conclusion Author Contribution Acknowledgement References

Kamruzzaman Mithu 1*

+ Author Affiliations

Data Modeling 1 (1) 1-11 https://doi.org/10.25163/data.1110858

Submitted: 25 September 2020 Revised: 12 November 2020  Accepted: 23 November 2020  Published: 25 November 2020 


Abstract

Background. Prediabetes affects an estimated 96 million American adults, yet more than 80% remain unaware of their condition — a gap that matters, since prediabetes is largely reversible while type 2 diabetes generally is not. Despite this, most predictive modeling efforts in the diabetes literature have focused on undiagnosed diabetes broadly, leaving prediabetes-specific prediction comparatively underexplored, partly because of scarce labeled data and persistent class imbalance. Methods. This study used a retrospective dataset of 1,000 patient records sourced from Medical City Hospital and the Al-Kindy Teaching Hospital in Iraq, comprising demographic and routine laboratory variables (e.g., BMI, HbA1c, lipid profile, blood urea). Feature selection combined manual review with information-gain analysis; outliers were identified and removed using boxplot- and histogram-based screening; and class imbalance across non-diabetic, prediabetic, and diabetic groups was addressed using SMOTE-based oversampling. Two classifiers — K-Nearest Neighbors and Decision Tree — were trained on a 70/30 split and evaluated using accuracy, precision, recall, F1-score, and AUROC. Results. The Decision Tree model achieved 82% accuracy, modestly outperforming KNN's 79%. Both models classified diabetic patients more reliably than prediabetic ones, a pattern attributable largely to the underlying class imbalance that oversampling only partially corrected. Conclusion. Reasonably accurate prediabetes screening appears achievable using a small set of accessible, low-cost variables, though prediabetes-specific detection remains disproportionately difficult given current data limitations. Expanding future datasets to include lifestyle and family-history variables, alongside testing ensemble methods, may meaningfully improve model reliability for this often-overlooked, high-risk population.

Keywords: prediabetes; machine learning; K-nearest neighbors; decision tree; early detection

References

American Diabetes Association. (2019). Prediabetes. https://diabetes.org/diabetes/prediabetes

Centers for Disease Control and Prevention. (2019). Prediabetes: Your chance to prevent type 2 diabetes. https://www.cdc.gov/diabetes/basics/prediabetes.html

Centers for Disease Control and Prevention. (2018). National diabetes statistics report. https://www.cdc.gov/diabetes/data/statistics-report/index.html

Choi, S. B., Kim, W. J., Yoo, T. K., Park, J. S., Chung, J. W., Lee, Y., Kang, E. S., & Kim, D. W. (2014). Screening for prediabetes using machine learning models. Computational and Mathematical Methods in Medicine, 2014, Article 618976. https://doi.org/10.1155/2014/618976

Deberneh, H. M., & Kim, I. (2021). Prediction of type 2 diabetes based on machine learning algorithm. International Journal of Environmental Research and Public Health, 18(6), Article 3317. https://doi.org/10.3390/ijerph18063317

Erlin, Junadhi, Agustin, L., Fajri, N., & Sari, S. N. (2020). Early detection of diabetes using machine learning with logistic regression algorithm. Jurnal Nasional Teknik Elektro dan Teknologi Informasi. https://jurnal.ugm.ac.id/v3/JNTETI/article/view/3586/1647

Hasan, M. K., Alam, M. A., Das, D., Hossain, E., & Hasan, M. (2020). Diabetes prediction using ensembling of different machine learning classifiers. IEEE Access, 8, 76516-76531. https://doi.org/10.1109/ACCESS.2020.2989857

International Diabetes Federation. (2020). IDF diabetes atlas (10th ed.). https://diabetesatlas.org/

Islam, R. M., Islam, S. M. S., Ferrari, A. J., Rahman, M. M., & Rahman, M. A. (2021). Prevalence of diabetes and prediabetes among Bangladeshi adults and associated factors: Evidence from the demographic and health survey, 2017-18. medRxiv. https://doi.org/10.1101/2021.01.26.21250519

Mahboob Alam, T., Iqbal, M. A., Ali, Y., Wahab, A., Ijaz, S., Baig, T. I., Hussain, A., Malik, M. A., Raza, M. M., Ibrar, S., & Abbas, Z. (2019). A model for early prediction of diabetes. Informatics in Medicine Unlocked, 16, Article 100204. https://doi.org/10.1016/j.imu.2019.100204

Mayo Clinic Staff. (2018). Prediabetes: Symptoms and causes. Mayo Clinic. https://www.mayoclinic.org/diseases-conditions/prediabetes/symptoms-causes/syc-20355278

MedlinePlus. (2018). Prediabetes. National Library of Medicine. https://medlineplus.gov/prediabetes.html

Rashid, A. (2020). Diabetes dataset [Data set]. Mendeley Data. https://data.mendeley.com/datasets/wj9rwkp9c2/1

Tigga, N. P., & Garg, S. (2020). Prediction of type 2 diabetes using machine learning classification methods. Procedia Computer Science, 167, 706-716. https://doi.org/10.1016/j.procs.2020.03.336


Article metrics
View details
1
Downloads
0
Citations
18
Views
📖 Cite article

View Dimensions


View Plumx


View Altmetric



1
Save
0
Citation
18
View
0
Share