Ultra-High Dimensional Feature Selection and Mean Estimation under Missing at Random
- 1 Department of Basic Sciences, Guilin University of Technology at Nanning, Chongzuo, China
- 2 College of Science, Guilin University of Technology, Guilin, China
- 3 Department of Basic Sciences, Guilin University of Technology at Nanning, Chongzuo, China
Abstract
Next Generation Sequencing (NGS) provides an effective basis for estimating the survival time of cancer patients, but it also poses the problem of high data dimensionality, in addition to the fact that some patients drop out of the study, making the data missing, so a method for estimating the mean of the response variable with missing values for the ultra-high dimensional datasets is needed. In this paper, we propose a two-stage ultra-high dimensional variable screening method, RF-SIS, based on random forest regression, which effectively solves the problem of estimating missing values due to excessive data dimension. After the dimension reduction process by applying RF-SIS, mean interpolation is executed on the missing responses. The results of the simulated data show that compared with the estimation method of directly deleting missing observations, the estimation results of RF-SIS-MI have significant advantages in terms of the proportion of intervals covered, the average length of intervals, and the average absolute deviation.
- Barzi, F. and Woodward, M. (2004) Imputations of Missing Values in Practice: Results from Imputations of Serum Cholesterol in 28 Cohort Studies. American Journal of Epidemiology, 160, 34-45. https://doi.org/10.1093/aje/kwh175
- Tibshirani, R. (1996) Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society: Series B (Methodological), 58, 267-288. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x
- Fan, J.Q. and Lv, J.C. (2008) Sure Independence Screening for Ultrahigh Dimensional Feature Space. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 70, 849-911. https://doi.org/10.1111/j.1467-9868.2008.00674.x
- Fan, J. and Song, R. (2010) Sure Independence Screening in Generalized Linear Models with NP-Dimensionality. The Annals of Statistics, 38, 3567-3604. https://doi.org/10.1214/10-AOS798
- Li, K., Wang, F., Yang, L. and Liu, R. (2023) Deep Feature Screening: Feature Selection for Ultra-High-Dimensional Data via Deep Neural Networks. Neurocomputing, 538, Article ID: 126186. https://doi.org/10.1016/j.neucom.2023.03.047
- Zhou, L. and Wang, H. (2022) A Combined Feature Screening Approach of Random Forest and Filterbased Methods for Ultra-High Dimensional Data. Current Bioinformatics, 17, 344-357. https://doi.org/10.2174/1574893617666220221120618
- Cheng, X. and Wang, H. (2022) A Generic Model-Free Feature Screening Procedure for Ultra-High Dimensional Data with Categorical Response. Computer Methods and Programs in Biomedicine, 229, Article ID: 107269. https://doi.org/10.1016/j.cmpb.2022.107269
- Zou, H. and Hastie, T. (2005) Regularization and Variable Selection via the Elastic Net. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 67, 301-320. https://doi.org/10.1111/j.1467-9868.2005.00503.x
- Zou, H. (2006) The Adaptive Lasso and Its Oracle Properties. Journal of the American Statistical Association, 101, 1418-1429. https://doi.org/10.1198/016214506000000735
- Fan, J. and Li, R. (2001) Variable Selection via Nonconcave Penalized Likelihood and Its Oracle Properties. Journal of the American Statistical Association, 96, 1348-1360. https://doi.org/10.1198/016214501753382273
- Breiman, L. (2001) Random Forests. Machine Learning, 45, 5-32. https://doi.org/10.1023/A:1010933404324
- Little, R.J. and Rubin, D.B. (2019) Statistical Analysis with Missing Data. John Wiley and Sons, Hoboken, NJ. https://doi.org/10.1002/9781119482260