Estimating the Empirical Null Distribution of Maxmean Statistics in Gene Set Analysis
- 1 Department of Biostatistics, University at Buffalo, Buffalo, USA
- 2 Department of Biostatistics and Bioinformatics, Roswell Park Cancer Institute, Buffalo, USA
- 3 Department of Biostatistics and Bioinformatics, Roswell Park Cancer Institute, Buffalo, USA
- 4 Department of Biostatistics, University at Buffalo, Buffalo, USA
Abstract
Gene Set Analysis (GSA) is a framework for testing the association of a set of genes and the outcome, e.g. disease status or treatment group. The method replies on computing a maxmean statistic and estimating the null distribution of the maxmean statistics via a restandardization procedure. In practice, the pre-determined gene sets have stronger intra-correlation than genes across sets. This may result in biases in the estimated null distribution. We derive an asymptotic null distribution of the maxmean statistics based on sparsity assumption. We propose a flexible two group mixture model for the maxmean statistics. The mixture model allows us to estimate the null parameters empirically via maximum likelihood approach. Our empirical method is compared with the restandardization procedure of GSA in simulations. We show that our method is more accurate in null density estimation when the genes are strongly correlated within gene sets.
- Nam, D. and Kim, S.-Y. (2008) Gene-Set Approach for Expression Pattern Analysis. Briefings in Bioinformatics, 9, 189-197. https://doi.org/10.1093/bib/bbn001
- Subramanian, A., Tamayo, P., Mootha, V.K., Mukherjee, S., Ebert, B.L., Gillette, M.A., Paulovich, A., Pomeroy, S.L., Golub, T.R., Lander, E.S., et al. (2005) Gene Set Enrichment Analysis: A Knowledge-Based Approach for Interpreting Genome-Wide Expression Profiles. Proceedings of the National Academy of Sciences of the United States of America, 102, 15545-15550. https://doi.org/10.1073/pnas.0506580102
- Efron, B. and Tibshirani, R. (2007) On Testing the Significance of Sets of Genes. The Annals of Applied Statistics, 1, 107-129. https://doi.org/10.1214/07-AOAS101
- Kanehisa, M. and Goto, S. (2000) KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research, 28, 27-30. https://doi.org/10.1093/nar/28.1.27
- Ashburner, M., Ball, C.A., Blake, J.A., Botstein, D., Butler, H., Cherry, J.M., Davis, A.P., Dolinski, K., Dwight, S.S., Eppig, J.T., et al. (2000) Gene Ontology: Tool for the Unification of Biology. Nature Genetics, 25, 25-29. https://doi.org/10.1038/75556
- Basu, A. and Ghosh, J. (1978) Identifiability of the Multinormal and Other Distributions under Competing Risks Model. Journal of Multivariate Analysis, 8, 413-429. https://doi.org/10.1016/0047-259X(78)90064-7
- Cain, M. (1994) The Moment-Generating Function of the Minimum of Bivariate Normal Random Variables. The American Statistician, 48, 124-125.
- Billingsley, P. (1995) Probability and Measure. 3rd Edition, Wiley Series in Probability and Mathematical Statistics.
- Efron, B. (2004) Large-Scale Simultaneous Hypothesis Testing. Journal of the American Statistical Association, 99, 96-104. https://doi.org/10.1198/016214504000000089
- Efron, B. (2012) Large-Scale Inference: Empirical Bayes Methods for Estimation, Testing, and Prediction, Volume 1. Cambridge University Press.
- Langfelder, P. and Horvath, S. (2008) WGCNA: An R Package for Weighted Correlation Network Analysis. BMC Bioinformatics, 9, 1-13. https://doi.org/10.1186/1471-2105-9-559
- Benjamini, Y. and Hochberg, Y. (1995) Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society. Series B (Methodological), 57, 289-300.