An Empirical Study of Downstream Analysis Effects of Model Pre-Processing Choices
- 1 Analytics and Data Science Institute, College of Software and Computer Engineering, Kennesaw State University, Kennesaw, GA, USA
- 2 Analytics and Data Science Institute, College of Software and Computer Engineering, Kennesaw State University, Kennesaw, GA, USA
Abstract
This study uses an empirical analysis to quantify the downstream analysis effects of data pre-processing choices. Bootstrap data simulation is used to measure the bias-variance decomposition of an empirical risk function, mean square error (MSE). Results of the risk function decomposition are used to measure the effects of model development choices on model bias, variance, and irreducible error. Measurements of bias and variance are then applied as diagnostic procedures for model pre-processing and development. Best performing model-normalization-data structure combinations were found to illustrate the downstream analysis effects of these model development choices. In addition s , results found from simulations were verified and expanded to include additional data characteristics (imbalanced, sparse) by testing on benchmark datasets available from the UCI Machine Learning Library. Normalization results on benchmark data were consistent with those found using simulations, while also illustrating that more complex and/or non-linear models provide better performance on datasets with additional complexities. Finally, applying the findings from simulation experiments to previously tested applications led to equivalent or improved results with less model development overhead and processing time.
- Wolpert, D.H., Macready, W.G., et al. (1997) No Free Lunch Theorems for Optimization. IEEE Transactions on Evolutionary Computation, 1, 67-82. https://doi.org/10.1109/4235.585893
- Carp, J. (2012) On the Plurality of (Methodological) Worlds: Estimating the Analytic Flexibility of fMRI Experiments. Frontiers in Neuroscience, 6, 149. https://doi.org/10.3389/fnins.2012.00149
- Carp, J. (2012) The Secret Lives of Experiments: Methods Reporting in the fMRI Literature. Neuroimage, 63, 289-300. https://doi.org/10.1016/j.neuroimage.2012.07.004
- Wagenmakers, E.-J., et al. (2012) An Agenda for Purely Confirmatory Research. Perspectives on Psychological Science, 7, 632-638. https://doi.org/10.1177/1745691612463078
- Botvinik-Nezer, R., et al. (2020) Variability in the Analysis of a Single Neuroimaging Dataset by Many Teams. Nature, 1-7.
- Silberzahn, R., et al. (2018) Many Analysts, One Data Set: Making Transparent How Variations in Analytic Choices Affect Results. Advances in Methods and Practices in Psychological Science, 1, 337-356.
- Yuan, L.H., et al. (2015) A Mixture-of-Modelers Approach to Forecasting NCAA Tournament Outcomes. Journal of Quantitative Analysis in Sports, 11, 13-27. https://doi.org/10.1515/jqas-2014-0056
- Singh, S. (2018) Understanding the Bias-Variance Tradeoff. https://towardsdatascience.com/understanding-the-bias-variance-tradeoff-165e6942b229
- Domingos, P. (2000) A Unified Bias-Variance Decomposition. Proceedings of 17th International Conference on Machine Learning, 231-238.
- Dietterich, T.G. and Kong, E.B. (1995) Machine Learning Bias, Statistical Bias, and Statistical Variance of Decision Tree Algorithms. Technical Report, Department of Computer Science, Oregon State University, Corvallis.
- Normalization. https://www.codecademy.com/articles/normalization
- Evans, C., Hardin, J. and Stoebel, D.M. (2017) Selecting Between-Sample RNA-Seq Normalization Methods from the Perspective of Their Assumptions. Briefings in Bioinformatics, 19, 776-792. https://doi.org/10.1093/bib/bbx008
- Agresti, A. (2003) Categorical Data Analysis. Vol. 482, John Wiley & Sons, Hoboken. https://doi.org/10.1002/0471249688
- Owen, S., et al. (2015) Advanced Analytics with Apache Spark.
- Chakure, A. (2020) Random Forest and Its Implementation. https://towardsdatascience.com/random-forest-and-its-implementation-71824ced454f