NSU-XGBoost: A Leakage-Controlled XGBoost Framework with Threshold Optimization for Improved Disease Detection
- 1 Department of Nursing and Allied Health, Norfolk State University, Norfolk, USA
Abstract
Accurate early detection of diabetes is an important objective in clinical decision support and population health analytics. Machine-learning models can improve risk classification; however, apparent gains in performance may be misleading when models are evaluated using different test populations or when information from the test set influences model development. This study presents NSU-XGBoost , a leakage-controlled, SMOTE-enhanced, and hyperparameter-optimized XGBoost framework for diabetes detection. The original dataset of 768 observations was divided into a training cohort of 614 observations and a frozen independent test cohort of 154 observations. The test set remained unchanged and was excluded from all model development, augmentation, and optimization procedures. Synthetic Minority Oversampling Technique (SMOTE) was confined to the training workflow to improve minority-class learning without introducing synthetic observations into the independent evaluation cohort. XGBoost hyperparameters were optimized using five-fold stratified cross-validation within the training data. On the frozen test cohort, tuned XGBoost without augmentation achieved 72.1% accuracy, 51.9% sensitivity, 83.0% specificity, an F1-score of 56.6%, ROC-AUC of 0.818, and PR-AUC of 0.661. NSU-XGBoost improved accuracy to 74.7%, sensitivity to 81.5%, F1-score to 69.3%, ROC-AUC to 0.835, and PR-AUC to 0.725. False-negative classifications decreased from 26 to 10, although specificity declined from 83.0% to 71.0%, demonstrating the expected tradeoff between improved diabetes detection and increased false-positive classifications. The close agreement between cross-validated ROC-AUC and independent-test ROC-AUC indicated no evidence of severe augmentation-induced overfitting. These findings demonstrate that NSU-XGBoost can substantially improve the detection of diabetes-positive cases under a strictly leakage-controlled evaluation framework. The proposed approach emphasizes independent validation, training-only augmentation, and clinically meaningful error analysis rather than reliance on overall accuracy alone.
- De Melo, P., and St. Rose, M. (2025) Accurate Classification of Diabetes via PM Generative AI. Advances in Bioscience and Biotechnology , 16, 379-409. https://doi.org/10.4236/abb.2025.169025
- Kadhm, M.S., Ghindawi, I.W. and Mhawi, D.E. (2018) An Accurate Diabetes Prediction System Based on K-Means Clustering and Proposed Classification Approach. International Journal of Applied Engineering Research , 13, 4038-4041.
- Sneha, N. and Gangil, T. (2019) Analysis of Diabetes Mellitus for Early Prediction Using Optimal Features Selection. Journal of Big Data , 6, Article No. 13. https://doi.org/10.1186/s40537-019-0175-6
- Zhu, C., Idemudia, C.U. and Feng, W. (2019) Improved Logistic Regression Model for Diabetes Prediction by Integrating PCA and K-Means Techniques. Informatics in Medicine Unlocked , 17, Article 100179. https://doi.org/10.1016/j.imu.2019.100179
- Sinaga, K.P. and Yang, M. (2020) Unsupervised K-Means Clustering Algorithm. IEEE Access , 8, 80716-80727. https://doi.org/10.1109/access.2020.2988796
- Wee, B.F., Sivakumar, S., Lim, K.H., Wong, W.K. and Juwono, F.H. (2023) Diabetes Detection Based on Machine Learning and Deep Learning Approaches. Multimedia Tools and Applications , 83, 24153-24185. https://doi.org/10.1007/s11042-023-16407-5
- Shakeel, P.M., Baskar, S., Dhulipala, V.R.S. and Jaber, M.M. (2018) Cloud Based Framework for Diagnosis of Diabetes Mellitus Using K-Means Clustering. Health Information Science and Systems , 6, Article No. 16. https://doi.org/10.1007/s13755-018-0054-0
- Ramadhan, N.G., Adiwijaya and Romadhony, A. (2021) Preprocessing Handling to Enhance Detection of Type 2 Diabetes Mellitus Based on Random Forest. International Journal of Advanced Computer Science and Applications , 12, 223-228. https://doi.org/10.14569/ijacsa.2021.0120726
- Saxena, S., Mohapatra, D., Padhee, S. and Sahoo, G.K. (2023) Machine Learning Algorithms for Diabetes Detection: A Comparative Evaluation of Performance of Algorithms. Evolutionary Intelligence , 16, 587-603. https://doi.org/10.1007/s12065-021-00685-9
- Wang, J. and Su, X. (2011) An Improved K-Means Clustering Algorithm. 2011 IEEE 3 rd International Conference on Communication Software and Networks , Xi’an, 27-29 May 2011, 44-46. https://doi.org/10.1109/iccsn.2011.6014384
- de Melo, P. and Davtyan, M. (2023) High Accuracy Classification of Populations with Breast Cancer: SVM Approach. Cancer Research Journal , 11, 94-104. https://doi.org/10.11648/j.crj.20231103.13