Comparison of Several Data Mining Methods in Credit Card Default Prediction
- 1 School of Science, Guilin University of Technology, Guilin, China
- 2 School of Science, Guilin University of Technology, Guilin, China
Abstract
LightGBM is an open-source, distributed and high-performance GB framework built by Microsoft company. LightGBM has some advantages such as fast learning speed, high parallelism efficiency and high-volume data, and so on. Based on the open data set of credit card in Taiwan region, five data mining methods, Logistic regression, SVM, neural network, Xgboost and LightGBM, are compared in this paper. The results show that the AUC, F 1 -Score and the predictive correct ratio of LightGBM are the best, and that of Xgboost is second. It indicates that LightGBM or Xgboost has a good performance in the prediction of categorical response variables and has a good application value in the big data era.
- Lyn, C. and Thomas, A. (2000) Survey of Credit and Behavioral Scoring: Forecasting Financial Risk of Lending to Consumers. International Journal of Forecasting, 16, 149-172. https://doi.org/10.1016/S0169-2070(00)00034-0
- Yeh, I.-C. and Lien, C.-H. (2009) The Comparisons of Data Mining Techniques for the Predictive Accuracy of Probability of Default of Credit Card Clients. Expert Systems with Applications, 36, 2473-2480. https://doi.org/10.1016/j.eswa.2007.12.020
- Mei, R.T., Xu, Y. and Wang, G.C. (2016) Analysis of Credit Card Default Prediction Model and Its Influencing Factors. Statistics and Application, 5, 263-275.
- Ye, Q.Y., Rao, H. and Ji, M.S. (2017) Business Sales Forecast Based on XGBOOST. Journal of Nanchang University, No. 3, 275-281.
- Guo, L.K., Qi, M., Finley, T., Wang, T.F., Chen, W., Ma, W.D., Ye, Q.W. and Liu, T.-Y. (2017) LightGBM: A Highly Efficient Gradient Boosting Decision Tree. 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4-9 December 2017, 1-9.
- Powers, D.M.W. (2011) Evaluation: From Precision, Recall and F-Measure to Roc, Informedness, Markedness & Correlation. Journal of Machine Learning Technologies, 2, 37-63.