Comparative Performance of Propensity Score Methods for Clinical Multi-Group Data: Balancing Confounders and Estimating Treatment Effects
- 1 Department of Biostatistics, School of Public Health, Southern Medical University, Guangzhou, China
- 2 Department of Biostatistics, School of Public Health, Southern Medical University, Guangzhou, China
- 3 Department of Biostatistics, School of Public Health, Southern Medical University, Guangzhou, China
- 4 Department of Biostatistics, School of Public Health, Southern Medical University, Guangzhou, China
- 5 Department of Biostatistics, School of Public Health, Southern Medical University, Guangzhou, China
- 6 School of International Education, Southern Medical University, Guangzhou, China
- 7 School of International Education, Southern Medical University, Guangzhou, China
- 8 Guangdong Experimental High School, Guangzhou, China
Abstract
Background: Propensity score methods have become a cornerstone of modern causal inference, enabling researchers to approximate the conditions of randomized experiments in observational studies. Despite their widespread adoption, most established propensity score approaches were originally developed for two-group comparisons, leaving a notable methodological gap for multi-group data commonly encountered in clinical trials, public health interventions, and comparative effectiveness research. Methods: We conducted a comparative evaluation of several propensity score methods in balancing confounders and estimating treatment effects using Monte Carlo simulation. Datasets of varying sample sizes were generated under two distinct hybrid data-generating structures. Propensity scores were estimated using both generalized linear models (GLM) and generalized boosting models (GBM), and were subsequently applied via inverse probability of treatment weighting (IPTW), overlap weighting (OW), and matching. Five specific method combinations were evaluated: GLM-IPTW, GLM-OW, GLM-matching, GBM-IPTW, and GBM-OW. Covariate balance was assessed using standardized mean differences (SMD), while treatment effect estimation performance was evaluated based on point estimate accuracy and root mean square error (RMSE). Results: Across simulation scenarios with both linear and non-linear underlying relationships, the GLM-matching approach generally outperformed other methods. GLM-OW and GBM-OW demonstrated superior performance in achieving covariate balance, while GLM-IPTW and GBM-IPTW yielded more accurate point estimates of the treatment effect. Conclusion: When the relationship between covariates and outcome is relatively simple and treatment assignment follows a linear model, the GLM-matching method proved particularly advantageous. It produced estimates closer to the true value and exhibited a stronger ability to balance covariates compared to the other methods considered.
- Kendall, M.A., Zander, T., Wolansky, R.L., Teixeira, L. and Kuo, P.C. (2025) Propensity Score Matching: A Step-by-Step Guide to Coding in R and Application in Observational Research Studies. The American Surgeon™ , 91, 1949-1955. https://doi.org/10.1177/00031348251331293
- Shurrab, M., Ko, D.T., Jackevicius, C.A., Tu, K., Middleton, A., Michael, F., et al . (2023) A Review of the Use of Propensity Score Methods with Multiple Treatment Groups in the General Internal Medicine Literature. Pharmacoepidemiology and Drug Safety , 32, 817-831. https://doi.org/10.1002/pds.5635
- McCaffrey, D.F., Griffin, B.A., Almirall, D., Slaughter, M.E., Ramchand, R. and Burgette, L.F. (2013) A Tutorial on Propensity Score Estimation for Multiple Treatments Using Generalized Boosted Models. Statistics in Medicine , 32, 3388-3414. https://doi.org/10.1002/sim.5753
- Wang, J. and Marion-Gallois, R. (2022) Propensity Score Matching and Stratification Using Multiparty Data without Pooling. Pharmaceutical Statistics , 22, 4-19. https://doi.org/10.1002/pst.2250
- Yang, S., Zhou, R., Li, F. and Thomas, L.E. (2023) Propensity Score Weighting Methods for Causal Subgroup Analysis with Time-to-Event Outcomes. Statistical Methods in Medical Research , 32, 1919-1935. https://doi.org/10.1177/09622802231188517
- Woo, M., Reiter, J.P. and Karr, A.F. (2008) Estimation of Propensity Scores Using Generalized Additive Models. Statistics in Medicine , 27, 3805-3816. https://doi.org/10.1002/sim.3278
- Gabriel, E.E., Sachs, M.C., Martinussen, T., Waernbaum, I., Goetghebeur, E., Vansteelandt, S., et al . (2023) Inverse Probability of Treatment Weighting with Generalized Linear Outcome Models for Doubly Robust Estimation. Statistics in Medicine , 43, 534-547. https://doi.org/10.1002/sim.9969
- Judkins, D.R. and Porter, K.E. (2015) Robustness of Ordinary Least Squares in Randomized Clinical Trials. Statistics in Medicine , 35, 1763-1773. https://doi.org/10.1002/sim.6839
- Tang, T., Austin, P.C., Lawson, K.A., Finelli, A. and Saarela, O. (2020) Constructing Inverse Probability Weights for Institutional Comparisons in Healthcare. Statistics in Medicine , 39, 3156-3172. https://doi.org/10.1002/sim.8657
- Li, L. and Greene, T. (2013) A Weighting Analogue to Pair Matching in Propensity Score Analysis. The International Journal of Biostatistics , 9, 215-234. https://doi.org/10.1515/ijb-2012-0030
- Mlcoch, T., Hrnciarova, T., Tuzil, J., Zadak, J., Marian, M. and Dolezal, T. (2019) Propensity Score Weighting Using Overlap Weights: A New Method Applied to Regorafenib Clinical Data and a Cost-Effectiveness Analysis. Value in Health , 22, 1370-1377. https://doi.org/10.1016/j.jval.2019.06.010