A Hybrid Machine Learning Framework for Early Diabetes Prediction in Sierra Leone Using Feature Selection and Soft-Voting Ensemble
- 1 College of Software, Nankai University, Tianjin, China
Abstract
This paper proposes a hybrid machine learning framework for early diabetes prediction tailored to Sierra Leone, where locally representative datasets are scarce. The framework integrates Random Forest (RF), Logistic Regression (LR), and Extreme Gradient Boosting (XGBoost) into a probability-based soft-voting ensemble that prioritizes sensitivity (recall) for screening. Experiments were conducted under two conditions: 1) using all available features and 2) after feature selection based on RF importance, retaining six clinically meaningful predictors (Glucose, Diabetes Pedigree Function, Skin Thickness, Age, Body Mass Index, and Insulin). Evaluation employed Accuracy, Precision, Recall, F1-score, ROC-AUC, and confusion-matrix analysis with a screening-oriented decision threshold. Before feature selection, the hybrid model achieved a Recall of 0.8571 and an ROC-AUC of 0.8610, reducing false negatives compared with individual classifiers. After feature selection, performance remained competitive while improving interpretability and deployment feasibility. Benchmark validation on the Pima Indians Diabetes dataset further supported the robustness of the approach. The proposed hybrid framework provides a practical, sensitivity-focused decision-support tool for early diabetes screening in low-resource clinical environments.
- Magliano, D.J., Boyko, E.J., IDF Diabetes Atlas, et al . (2021) Covid-19 and Diabetes. In IDF Diabetes Atlas [Internet]. 10th Edition, International Diabetes Federation.
- World Health Organization, et al . (2020) World Health Statistics 2020.
- Atun, R., Davies, J.I., Gale, E.A.M., et al . (2017) Diabetes in Sub-Saharan Africa: From Clinical Care to Health Policy. The Lancet Diabetes & Endocrinology , 5, 622-667.
- World Health Organization, et al . (2017) Who Country Cooperation Strategy 2017-2021: Sierra Leone.
- Kaviyaadharshani, D., Nivedhidha, M., Jeyarohini, R., Rani, J.L.E., Ramkumar, M.P. and Selvan, G.S.R.E. (2024) Diagnosing Diabetes Using Machine Learning-Based Predictive Models. Procedia Computer Science , 233, 288-294. https://doi.org/10.1016/j.procs.2024.03.218
- Dutta, A., Hasan, M.K., Ahmad, M., Awal, M.A., Islam, M.A., Masud, M., et al . (2022) Early Prediction of Diabetes Using an Ensemble of Machine Learning Models. Inter national Journal of Environmental Research and Public Health , 19, Article No. 12378. https://doi.org/10.3390/ijerph191912378
- He, H.B. and Garcia, E.A. (2009) Learning from Imbalanced Data. IEEE Transactions on Knowledge and Data Engineering , 21, 1263-1284. https://doi.org/10.1109/tkde.2008.239
- Bachmann, K.N. and Wang, T.J. (2017) Biomarkers of Cardiovascular Disease: Contributions to Risk Prediction in Individuals with Diabetes. Diabetologia , 61, 987-995. https://doi.org/10.1007/s00125-017-4442-9
- Hosmer, D.W., Lemeshow, S. and Sturdivant, R.X. (2013) Applied Logistic Regression. Wiley. https://doi.org/10.1002/9781118548387
- Breiman, L. (2001) Random Forests. Machine Learning , 45, 5-32. https://doi.org/10.1023/a:1010933404324
- Chen, T.Q. (2016) Xgboost: A Scalable Tree Boosting System. Cornell University.
- Fawcett, T. (2006) An Introduction to ROC Analysis. Pattern Recognition Letters , 27, 861-874. https://doi.org/10.1016/j.patrec.2005.10.010
- UCI Machine Learning (2016) Pima Indians Diabetes Database. https://kaggle.com/uciml/pima-indians-diabetes-database
- Smith, J.W., Everhart, J.E., Dickson, W.C., Knowler, W.C. and Johannes, R.S. (1988) Using the ADAP Learning Algorithm to Forecast the Onset of Diabetes Mellitus. Proceedi ngs of the Annual Symposium on Computer Application in Medical Care , Washington DC, 6-9 November 1988, 261.