A Predictive Modeling for Detecting Fraudulent Automobile Insurance Claims
- 1 Department of Mathematics and Statistics, California State University, Long Beach, CA, USA
- 2 Department of Mathematics and Statistics, California State University, Long Beach, CA, USA
- 3 Department of Mathematics and Statistics, California State University, Long Beach, CA, USA
Abstract
Fraudulent automobile insurance claims are not only a loss for insurance companies, but also for their policyholders. The goal of this research is to develop, first, a decision-making algorithm to classify whether a claim is classified as fraudulent or not; and, second, what types of variables should be focused to detect fraudulent claims. To achieve this goal, highly accurate prediction models are built by discovering important sets of features via variable selection algorithms, which can in turn help prevent future loss. In this research, parametric and nonparametric statistical learning algorithms are considered to reduce uncertainty and increase the chances of detecting the appropriate claims. An important set of features for a model is determined by measuring variable importance based on the observed characteristics of a claim via a cross-validation and by testing improvement of the performance at which automobile fraudulent claims are accurately classified using Akaike Information Criterion. We could achieve accuracy above 95% with a set of features selected via a cross-validation. This research would offer some benefit to the insurance industry for their fraud detection research in order to prevent insurance abuse from escalating any further.
- Viaene, S., Ayuso, M., Guillen, M., Gheel, D.V. and Dedene, G. (2007) Strategies for Detecting Fraudulent Claims in the Automobile Insurance Industry. European Journal of Operational Research, 176, 565-583. https://doi.org/10.1016/j.ejor.2005.08.005
- Boyer, M. (2000) Centralizing Insurance Fraud Investigation. The Geneva Papers on Risk and Insurance Theory, 25, 159-178. https://doi.org/10.1023/A:1008766413327
- Derrig, R.A. (2002) Insurance Fraud. Journal of Risk and Insurance, 69, 271-287. https://doi.org/10.1111/1539-6975.00026
- Ciaene, S., Derrig, R.A., Baesens, B. and Dedene, G. (2002) A Comparison of State-of-the-Art Classification Techniques for Expert Automobile Insurance Claim Fraud Detection. The Journal of Risk and Insurance, 69, 373-421. https://doi.org/10.1111/1539-6975.00023
- Caudill, S., Ayuso, M. and Guillen, M. (2005) Fraud Detection Using a Multinomial Logit Model with Missing Information. The Journal of Risk and Insurance, 72, 539-550. https://doi.org/10.1111/j.1539-6975.2005.00137.x
- Artis, M., Ayuso, M. and Cuillen, M. (2002) Detection of Automobile Insurance Fraud with Discrete Choice Models and Misclassified Claims. The Journal of Risk and Insurance, 69, 325-340. https://doi.org/10.1111/1539-6975.00022
- Wang, Y. and Xu, W. (2018) Leveraging Deep Learning with LDA-Based Text Analytics to Detect Automobile Insurance Fraud. Decision Support Systems, 105, 87-95. https://doi.org/10.1016/j.dss.2017.11.001
- Nian, K., Zhang, H., Tayal, A., Coleman, T. and Li, Y. (2016) Auto Insurance Fraud Detection Using Unsupervised Spectral Ranking for Anomaly. The Journal of Finance and Data Science, 2, 57-58. https://doi.org/10.1016/j.jfds.2016.03.001
- Tibshirani, R. (2011) Regression Shrinkage and Selection via the Lasso: A Retrospective. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73, 273-282. https://doi.org/10.1111/j.1467-9868.2011.00771.x
- Breiman, L. (2001) Random Forests. Machine Learning, 45, 5-32. https://doi.org/10.1023/A:1010933404324
- Cortes, C. and Vapnik, V. (1995) Support-Vector Networks. Machine Learning, 20, 273-297. https://doi.org/10.1007/BF00994018
- Pyle, D. (1999) Data Preparation for Data Mining. Morgan Kaufmann Publishers, Inc., San Francisco.
- Tan, P.-N., Steinbach, M. and Kumar, V. (2005) Introduction to Data Mining. Pearson Addison Wesley, Boston.