Research ArticleOpen AccessGoogle Scholar indexed
Analysis of the Resolution of Crime Using Predictive Modeling
Department of Statistics, Truman State University, Kirksville, MO, USA
Department of Physics, Truman State University, Kirksville, MO, USA
Department of Mathematics, SUNY Cortland, Cortland, NY, USA
Department of Mathematics, Northwood University, Midland, MI, USA
- 1 Department of Statistics, Truman State University, Kirksville, MO, USA
- 2 Department of Physics, Truman State University, Kirksville, MO, USA
- 3 Department of Mathematics, SUNY Cortland, Cortland, NY, USA
- 4 Department of Mathematics, Northwood University, Midland, MI, USA
Open Journal of Statistics·Volume 10 (2020)·Pages 600–610·Published 8 May 2020·DOI10.4236/ojs.2020.103036
Copy link · social · email
Abstract
There has been evidence of crime in the US since colonization. In this article, we analyze the crime statistics of San Francisco and its resolution of crime recorded from January to September of the year 2018. We define resolution of crime as a target variable and study its relationship with other variables. We make several classification models to predict resolution of crime using several data mining techniques and suggest the best model for predicting resolution.
KeywordsMachine LearningClassification Model ComparisonPredictive ModelingResolution of Crime
- 2017 Crime in the United States. https://ucr.fbi.gov/crime-in-the-u.s/2017/crime-in-the-u.s.-2017/topic-pages/property-crime
- Mitchell, T.M. (1997) Machine Learning. McGraw-Hill Higher Education, New York.
- Alpaydin, E. (2020) Introduction to Machine Learning. MIT Press, Cambridge.
- Bishop, C.M. (2006) Pattern Recognition and Machine Learning. Springer, Berlin.
- James, G., Witten, D., Hastie, T. and Tibshirani, R. (2013) An Introduction to Statistical Learning. Vol. 112, Springer, New York, 3-7. https://doi.org/10.1007/978-1-4614-7138-7
- Police Department Incident Report of City and County of San Francisco. https://data.sfgov.org/Public-Safety/Police-Department-Incident-Reports-2018-to-Present/wg3w-h783
- Guyon, I. and Elisseeff, A. (2003) An Introduction to Variable and Feature Selection. Journal of Machine Learning Research, 3, 1157-1182.
- Acuna, E. and Rodriguez, C. (2004) The Treatment of Missing Values and Its Effect on Classifier Accuracy. In: Classification, Clustering, and Data Mining Applications, Springer, Berlin, Heidelberg, 639-647. https://doi.org/10.1007/978-3-642-17103-1_60
- Van Buuren, S. (2018) Flexible Imputation of Missing Data. CRC Press, Boca Raton. https://doi.org/10.1201/9780429492259
- Crookston, N.L. and Finley, A.O. (2008) yaImpute: An R Package for kNN Imputation. Journal of Statistical Software, 23, 16 p. https://doi.org/10.18637/jss.v023.i10
- Cerda, P., Varoquaux, G. and Kégl, B. (2018) Similarity Encoding for Learning with Dirty Categorical Variables. Machine Learning, 107, 1477-1494. https://doi.org/10.1007/s10994-018-5724-2
- Asaithambi, S. and Why, H. (2017) Why, How and When to Scale Your Features. https://medium.com/greyatom/why-how-and-when-to-scale-your-features-4b30ab09db5e
- Manning, C. (2007) Logistic Regression (with R) Changes.
- Li, C. and Wang, B. (2014) Fisher Linear Discriminant Analysis. CCIS Northeastern University.
- Hansen, M., Dubayah, R. and DeFries, R. (1996) Classification Trees: An Alternative to Traditional Land Cover Classifiers. International Journal of Remote Sensing, 17, 1075-1081. https://doi.org/10.1080/01431169608949069
- Therneau, T., Atkinson, B., Ripley, B. and Ripley, M.B. (2015) Package “rpart”. http://cran.ma.ic.ac.uk/web/packages/rpart/rpart.pdf