TITLE:
Comparative Performance of Propensity Score Methods for Clinical Multi-Group Data: Balancing Confounders and Estimating Treatment Effects
AUTHORS:
Xinlan Teng, Jiayu Chen, Ying Guan, Kailiang Shen, Wenyi Lai, Emmanuel Nizeyimana, Teaway Angeline Patience, Eric Hu
KEYWORDS:
Multi-Group Data, Generalized Boosting Model, Generalized Linear Model, Overlap Weighting
JOURNAL NAME:
International Journal of Clinical Medicine,
Vol.17 No.1,
January
5,
2026
ABSTRACT: Background: Propensity score methods have become a cornerstone of modern causal inference, enabling researchers to approximate the conditions of randomized experiments in observational studies. Despite their widespread adoption, most established propensity score approaches were originally developed for two-group comparisons, leaving a notable methodological gap for multi-group data commonly encountered in clinical trials, public health interventions, and comparative effectiveness research. Methods: We conducted a comparative evaluation of several propensity score methods in balancing confounders and estimating treatment effects using Monte Carlo simulation. Datasets of varying sample sizes were generated under two distinct hybrid data-generating structures. Propensity scores were estimated using both generalized linear models (GLM) and generalized boosting models (GBM), and were subsequently applied via inverse probability of treatment weighting (IPTW), overlap weighting (OW), and matching. Five specific method combinations were evaluated: GLM-IPTW, GLM-OW, GLM-matching, GBM-IPTW, and GBM-OW. Covariate balance was assessed using standardized mean differences (SMD), while treatment effect estimation performance was evaluated based on point estimate accuracy and root mean square error (RMSE). Results: Across simulation scenarios with both linear and non-linear underlying relationships, the GLM-matching approach generally outperformed other methods. GLM-OW and GBM-OW demonstrated superior performance in achieving covariate balance, while GLM-IPTW and GBM-IPTW yielded more accurate point estimates of the treatment effect. Conclusion: When the relationship between covariates and outcome is relatively simple and treatment assignment follows a linear model, the GLM-matching method proved particularly advantageous. It produced estimates closer to the true value and exhibited a stronger ability to balance covariates compared to the other methods considered.