An Improved Treed Gaussian Process
- 1 Department of Statistics and Biostatistics, California State University, Hayward, CA, USA
- 2 Department of Statistics, University of California, Santa Cruz, CA, USA
Abstract
Many black box functions and datasets have regions of different variability. Models such as the Gaussian process may fall short in giving the best representation of these complex functions. One successful approach for modeling this type of nonstationarity is the Treed Gaussian process [1] , which extended the Gaussian process by dividing the input space into different regions using a binary tree algorithm. Each region became its own Gaussian process. This iterative inference process formed many different trees and thus, many different Gaussian processes. In the end these were combined to get a posterior predictive distribution at each point. The idea was that when the iterations were combined, smoothing would take place for the surface of the predicted points near tree boundaries. We introduce the Improved Treed Gaussian process, which divides the input space into a single main binary tree where the different tree regions have different variability. The parameters for the Gaussian process for each tree region are then determined. These parameters are then smoothed at the region boundaries. This smoothing leads to a set of parameters for each point in the input space that specify the covariance matrix used to predict the point. The advantage is that the prediction and actual errors are estimated better since the standard deviation and range parameters of each point are related to the variation of the region it is in. Further, smoothing between regions is better since each point prediction uses its parameters over the whole input space. Examples are given in this paper which show these advantages for lower-dimensional problems.
- Gramacy, R. and Lee, H.K. (2008) Bayesian Treed Gaussian Process Models with Application to Computer Modeling. Journal of the American Statistical Association, 103, 1119-1130. https://doi.org/10.1198/016214508000000689
- Santner, T., Williams, B. and Notz, W. (2003) The Design and Analysis of Computer Experiments. Springer, New York. https://doi.org/10.1007/978-1-4757-3799-8
- Kleijnen, J.P.C. (2015) Design and Analysis of Simulation Experiments. 2nd Edition, Springer, New York.
- Higdon, D., Swall, J. and Kern, J. (1999) Non-Stationary Spatial Modeling. Bayesian Statistics, 16, 761-768.
- Paciorek, C.J. and Schervish, M. (2006) Spatial Modelling Using a New Class of Nonstationary Covariance Functions. Environmentrics, 17, 483-506. https://doi.org/10.1002/env.785
- Liang, W.W.J. and Lee, H.K.H. (2019) Bayesian Nonstationary Gaussian Process Models for Large Datasets via Treed Process Convolutions. Advances in Data Analysis and Classification, 1-22.
- Kim, H., Mallick, B. and Holmes, C. (2005) Analyzing Nonstationary Spatial Data Using Piecewise Gaussian Processes. American Statistical Association, 100, 653-668. https://doi.org/10.1198/016214504000002014
- Heaton, M., Christensen, W. and Terres, M. (2014) Nonstationary Gaussian Process Models Using Spatial Hierarchical Clustering from Finite Differences. Technometrics, 59, 93-101. https://doi.org/10.1080/00401706.2015.1102763
- Li, F. and Sang, H. (2019) Spatial Homogeneity of Regression Coefficients for Large Datasets. American Statistical Association, 114, 1050-1062. https://doi.org/10.1080/01621459.2018.1529595
- Konomi, B., Sang, H. and Mallick, B. (2014) Adaptive Bayesian Nonstationary Modeling for Large Datasets Using Covariance Approximations. Computational and Graphical Statistics, 23, 802-829. https://doi.org/10.1080/10618600.2013.812872
- Kim, H., Mallick, B. and Holmes, C. (2005) Analyzing Nonstationary Spatial Data Using Piecewise Gaussian Processes. Journal of the American Statistical Association, 100, 653-668. https://doi.org/10.1198/016214504000002014
- Gramacy, R. (2005) Bayesian Treed Gaussian Process Models. PhD Thesis, UC Santa Cruz.