This paper presents a novel variable selection method in additive nonparametric regression model. This work is motivated by the need to select the number of nonparametric components and number of variables within each nonparametric component. The proposed method uses a combination of hard and soft shrinkages to separately control the number of additive components and the variables within each component. An efficient algorithm is developed to select the importance of variables and estimate the interaction network. Excellent performance is obtained in simulated and real data examples.
Manavalan, P. and Johnson, W.C. (1987) Variable Selection Method Improves the Prediction of Protein Secondary Structure from Circular Dichroism Spectra. Analytical Biochemistry, 167, 76-85.
Saeys, Y., Inza, I. and Larranaga, P. (2007) A Review of Feature Selection Techniques in Bioinformatics. Bioinformatics, 23, 2507-2517.
Kwon, O.-W., Chan, K., Hao, J. and Lee, T.-W. (2003) Emotion Recognition by Speech Signals. 8th European Conference on Speech Communication and Technology, Geneva, 1-4 September 2003.
Akaike, H. (1973) Maximum Likelihood Identification of Gaussian Autoregressive Moving Average Models. Biometrika, 60, 255-265. https://doi.org/10.1093/biomet/60.2.255
Schwarz, G., et al. (1978) Estimating the Dimension of a Model. The Annals of Statistics, 6, 461-464. https://doi.org/10.1214/aos/1176344136
Foster, D.P. and George, E.I. (1994) The Risk Ination Criterion for Multiple Regression. The Annals of Statistics, 22, 1947-1975.
Tibshirani, R. (1996) Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society. Series B (Methodological), 58, 267-288.
Efron, B., Hastie, T., Johnstone, I., Tibshirani, R., et al. (2004) Least Angle Regression. The Annals of Statistics, 32, 407-499. https://doi.org/10.1214/009053604000000067
Fan, J. and Li, R. (2001) Variable Selection via Nonconcave Penalized Likelihood and Its Oracle Properties. JASA, 96, 1348-1360. https://doi.org/10.1198/016214501753382273
Zou, H. and Hastie, T. (2005) Regularization and Variable Selection via the Elastic Net. Journal of the Royal Statistical Society, Series B, 67, 301-320. https://doi.org/10.1111/j.1467-9868.2005.00503.x
Zou, H. (2006) The Adaptive Lasso and Its Oracle Properties. JASA, 101, 1418-1429. https://doi.org/10.1198/016214506000000735
Zhang, C.-H. (2010) Nearly Unbiased Variable Selection under Minimax Concave Penalty. The Annals of Statistics, 38, 894-942. https://doi.org/10.1214/09-AOS729
Mitchell, T.J. and Beauchamp, J.J. (1988) Bayesian Variable Selection in Linear Regression. JASA, 83, 1023-1032. https://doi.org/10.1080/01621459.1988.10478694
George, E.I. and McCulloch, R.E. (1993) Variable Selection via Gibbs Sampling. Journal of the American Statistical Association, 88, 881-889. https://doi.org/10.1080/01621459.1993.10476353
George, E.I. and McCulloch, R.E. (1997) Approaches for Bayesian Variable Selection. Statistica sinica, 7, 339-373.
Laerty, J. and Wasserman, L. (2008) Rodeo: Sparse, Greedy Nonparametric Regression. The Annals of Statistics, 36, 28-63.
Wahba, G. (1990) Spline Models for Observational Data. Vol. 59, Siam. https://doi.org/10.1137/1.9781611970128
Green, P.J. and Silverman, B.W. (1993) Nonparametric Regression and Generalized Linear Models: A Roughness Penalty Approach. CRC Press, Boca Raton.
Hastie, T.J. and Tibshirani, R.J. (1990) Generalized Additive Models. Vol. 43, CRC Press, Boca Raton.
Lin, Y., Zhang, H.H., et al. (2006) Component Selection and Smoothing in Multivariate Nonparametric Regression. The Annals of Statistics, 34, 2272-2297. https://doi.org/10.1214/009053606000000722
Ravikumar, P., Laerty, J., Liu, H. and Wasserman, L. (2009) Sparse Additive Models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 71, 1009-1030. https://doi.org/10.1111/j.1467-9868.2009.00718.x
Radchenko, P. and James, G.M. (2010) Variable Selection Using Adaptive Nonlinear Interaction Structures in High Dimensions. Journal of the American Statistical Association, 105, 1541-1553. https://doi.org/10.1198/jasa.2010.tm10130
Liu, D., Lin, X. and Ghosh, D. (2007) Semiparametric Regression of Multidimensional Genetic Pathway Data: Least-Squares Kernel Machines and Linear Mixed Models. Biometrics, 63, 1079-1088. https://doi.org/10.1111/j.1541-0420.2007.00799.x
Zou, F., Huang, H., Lee, S. and Hoeschele, I. (2010) Nonparametric Bayesian Variable Selection with Applications to Multiple Quantitative Trait Loci Mapping with Epistasis and Gene-Environment Interaction. Genetics, 186, 385-394. https://doi.org/10.1534/genetics.109.113688
Savitsky, T., Vannucci, M. and Sha, N. (2011) Variable Selection for Nonparametric Gaussian Process Priors: Models and Computational Strategies. Statistical Science: A Review Journal of the Institute of Mathematical Statistics, 26, 130.
Yang, Y., Tokdar, S.T., et al. (2015) Minimax-Optimal Nonparametric Regression in High Dimensions. The Annals of Statistics, 43, 652-674. https://doi.org/10.1214/14-AOS1289
Fang, Z., Kim, I. and Schaumont, P. (2012) Flexible Variable Selection for Recovering Sparsity in Nonadditive Nonparametric Models.
Qamar, S. and Tokdar, S.T. (2014) Additive Gaussian Process Regression.
Rasmussen, C.E. and Williams, C.K.I. (2006) Gaussian Processes for Machine Learning.
Van der Vaart, A.W. and van Zanten, J.H. (2009) Adaptive Bayesian Estimation Using a Gaussian Random Field with Inverse Gamma Bandwidth. The Annals of Statistics, 37, 2655-2675.
Bhattacharya, A., Pati, D. and Dunson, D.B. (2014) Anisotropic Function Estimation Using Multi-Bandwidth Gaussian Processes. The Annals of Statistics, 42, 352-381. https://doi.org/10.1214/13-AOS1192
Tibshirani, R., et al. (1997) The Lasso Method for Variable Selection in the Cox Model. Statistics in Medicine, 16, 385-395. https://doi.org/10.1002/(SICI)1097-0258(19970228)16:4 3.0.CO;2-3
Hastie, T., Tibshirani, R., Friedman, J. and Franklin, J. (2005) The Elements of Statistical Learning: Data Mining, Inference and Prediction. The Mathematical Intelligencer, 27, 83-85. https://doi.org/10.1007/BF02985802
Park, T. and Casella, G. (2008) The Bayesian Lasso. Journal of the American Statistical Association, 103, 681-686. https://doi.org/10.1198/016214508000000337
Tipping, M.E. (2001) Sparse Bayesian Learning and the Relevance Vector Machine. The Journal of Machine Learning Research, 1, 211-244.
Grin, J.E. and Brown, P.J. (2010) Inference with Normal-Gamma Prior Distributions in Regression Problems. Bayesian Analysis, 5, 171-188. https://doi.org/10.1214/10-BA507
Carvalho, C.M., Polson, N.G. and Scott, J.G. (2010) The Horseshoe Estimator for Sparse Signals. Biometrika, 97, 465-480. https://doi.org/10.1093/biomet/asq017
Carvalho, C.M., Polson, N.G. and Scott, J.G. (2009) Handling Sparsity via the Horseshoe. International Conference on Artificial Intelligence and Statistics, Clearwater, 16-18 April 2009, 73-80.
Bhattacharya, A., Pati, D., Pillai, N.S. and Dunson, D.B. (2014) Dirichlet-Laplace Priors for Optimal Shrinkage. Journal of the American Statistical Association, 110, 1479-1490.
Polson, N.G. and Scott, J.G. (2010) Shrink Globally, Act Locally: Sparse Bayesian Regularization and Prediction. Bayesian Statistics, 9, 501-538.
Diebolt, J., Ip, E. and Olkin, I. (1994) A Stochastic EM Algorithm for Approximating the Maximum Likelihood Estimate. Technical Report 301, Department of Statistics, Stanford University, Stanford.
Meng, X.-L. and Rubin, D.B. (1994) On the Global and Component Wise Rates of Convergence of the EM Algorithm. Linear Algebra and Its Applications, 199, 413-425.
Hastings, W. (1970) Monte Carlo Sampling Methods Using Markov Chains and Their Applications. Biometrika, 57, 97-109. https://doi.org/10.1093/biomet/57.1.97
Bishop, C.M. (2006) Pattern Recognition and Machine Learning. Springer, Berlin.
Han, J., Kamber, M. and Pei, J. (2011) Data Mining: Concepts and Techniques: Concepts and Techniques. Elsevier, Amsterdam.
Guyon, I. and Elisseeff, A. (2003) An Introduction to Variable and Feature Selection. Journal of Machine Learning Research, 3, 1157-1182.
George Forman (2003) An Extensive Empirical Study of Feature Selection Metrics for Text Classification. Journal of Machine Learning Research, 3, 1289-1305.
Stoppiglia, H., Dreyfus, G., Dubois, R. and Oussar, Y. (2003) Ranking a Random Feature for Variable and Feature Selection. Journal of Machine Learning Research, 3, 1399-1414.
Genuer, R., Poggi, J.M. and Tuleau-Malot, C. (2010) Variable Selection Using Random Forests. Pattern Recognition Letters, 31, 2225-2236.
Chipman, H.A., George, E.I. and McCulloch, R.E. (2010) Bart: Bayesian Additive Regression Trees. The Annals of Applied Statistics, 4, 266-298.
Harrison, D. and Rubinfeld, D.L. (1978) Hedonic Housing Prices and the Demand for Clean Air. Journal of Environmental Economics and Management, 5, 81-102.
Yeh, I., et al. (2008) Modeling Slump of Concrete with Yash and Superplasticizer. Computers and Concrete, 5, 559-572. https://doi.org/10.12989/cac.2008.5.6.559
Yeh, I.C. (2007) Modeling Slumpow of Concrete Using Second-Order Regressions and Artificial Neural Networks. Cement and Concrete Composites, 29, 474-480.
Yurugi, M., Sakata, N., Iwai, M. and Sakai, G. (1993) Mix Proportion for Highly Workable Concrete. Proceedings of Concrete, Dundee, 7-9 September 1993, 579-589.
Redmond, M.A. and Highley, T. (2010) Empirical Analysis of Caseediting Approaches for Numeric Prediction. In: Innovations in Computing Sciences and Software Engineering, Springer, Berlin, 79-84.
Buczak, A.L. and Gifford, C.M. (2010) Fuzzy Association Rule Mining for Community Crime Pattern Discovery. In: ACM SIGKDD Workshop on Intelligence and Security Informatics, ACM, New York, 2.
Blake, C. and Merz, C.J. (1998) Repository of Machine Learning Databases.
Blumstein, A. and Rosenfeld, R. (2008) Factors Contributing to Us Crime Trends. In: Understanding Crime Trends: Workshop Report, The National Academies Press, Washington DC, 13-43.