Academic Journal
Choosing Covariate Balancing Methods for Causal Inference: Practical Insights From a Simulation Study.
| Title: | Choosing Covariate Balancing Methods for Causal Inference: Practical Insights From a Simulation Study. |
|---|---|
| Authors: | Peyrot E; Université Paris Cité and Université Sorbonne Paris Nord, Inserm, INRAE, Center for Research in Epidemiology and StatisticS (CRESS), Paris, France., Porcher R; Université Paris Cité and Université Sorbonne Paris Nord, Inserm, INRAE, Center for Research in Epidemiology and StatisticS (CRESS), Paris, France.; Centre d'Épidémiologie Clinique, Assistance Publique-Hôpitaux de Paris, Hôtel-Dieu, Paris, France., Petit F; Université Paris Cité and Université Sorbonne Paris Nord, Inserm, INRAE, Center for Research in Epidemiology and StatisticS (CRESS), Paris, France. |
| Source: | Statistics in medicine [Stat Med] 2026 Jul; Vol. 45 (15-17), pp. e70672. |
| Publication Type: | Journal Article |
| Language: | English |
| Journal Info: | Publisher: Wiley Country of Publication: England NLM ID: 8215016 Publication Model: Print Cited Medium: Internet ISSN: 1097-0258 (Electronic) Linking ISSN: 02776715 NLM ISO Abbreviation: Stat Med Subsets: MEDLINE |
| Imprint Name(s): | Original Publication: Chichester ; New York : Wiley, c1982- |
| MeSH Terms: | Observational Studies as Topic*/methods , Causality*, Monte Carlo Method ; Humans ; Computer Simulation ; Propensity Score ; Models, Statistical ; Confounding Factors, Epidemiologic ; Sample Size ; Least-Squares Analysis |
| Abstract: | Background: Weighting methods are widely used for confounding adjustment in observational studies, but their finite-sample behavior depends on implementation choices and empirical overlap. We compare IPTW, just- and over-identified covariate balancing propensity score (CBPS), CBPS by tailored-loss function (CBPS-TLF), energy balancing (EB), and kernel optimal matching (KOM). Methods: We conducted Monte Carlo simulations across 36 main scenarios varying sample size, treatment prevalence, and a complexity factor increasing confounding and reducing overlap. The main simulation considered a null constant treatment effect, with non-null constant effects examined as sensitivity analyses. Average treatment effects and average treatment effects on the treated were estimated using weighted least squares (WLS) and doubly robust (DR) estimators. Inference followed published recommendations when feasible. An empirical illustration used PROBITsim. Results: Performance depended on the estimator and scenario complexity. Under WLS, IPTW and CBPS-TLF were more sensitive to complexity, while standard CBPS often behaved similarly to IPTW but with less deterioration in some high-prevalence settings. EB and KOM showed more stable point-estimation patterns across scenarios. DR estimation reduced differences between weighting methods when all confounders were included in the outcome model, although confidence-interval performance remained heterogeneous. PROBITsim results were consistent with simulation patterns. Conclusions: The study should be read as practical guidance rather than a ranking of methods. It identifies settings where weighting analyses become sensitive to prevalence, overlap, tuning, and variance estimation. Confidence intervals that account for weight construction and tuning remain an important open practical issue. (© 2026 The Author(s). Statistics in Medicine published by John Wiley & Sons Ltd.) |
| References: | P. R. Rosenbaum and D. B. Rubin, “The Central Role of the Propensity Score in Observational Studies for Causal Effects,” Biometrika 70, no. 1 (1983): 41–55. J. K. Lunceford and M. Davidian, “Stratification and Weighting via the Propensity Score in Estimation of Causal Treatment Effects: A Comparative Study,” Statistics in Medicine 23, no. 19 (2004): 2937–2960. J. Hahn, “On the Role of the Propensity Score in Efficient Semiparametric Estimation of Average Treatment Effects,” Econometrica 66, no. 2 (1998): 315–331. J. M. Smit, J. H. Krijthe, K. WMR, et al., “Causal Inference Using Observational Intensive Care Unit Data: A Scoping Review and Recommendations for Future Practice,” NPJ Digital Medicine 6, no. 1 (2023): 221. H. Zuo, L. Yu, S. M. Campbell, S. S. Yamamoto, and Y. Yuan, “The Implementation of Target Trial Emulation for Causal Inference: A Scoping Review,” Journal of Clinical Epidemiology 162 (2023): 29–37. P. R. Rosenbaum, “Model‐Based Direct Adjustment,” Journal of the American Statistical Association 82, no. 398 (1987): 387–394. D. G. Horvitz and D. J. Thompson, “A Generalization of Sampling Without Replacement From a Finite Universe,” Journal of the American Statistical Association 47, no. 260 (1952): 663–685. Y. Xiao, E. E. M. Moodie, and M. Abrahamowicz, “Comparison of Approaches to Weight Truncation for Marginal Structural Cox Models,” Epidemiological Methods 2, no. 1 (2013): 1–20. T. Stürmer, M. Webster‐Clark, J. L. Lund, et al., “Propensity Score Weighting and Trimming Strategies for Reducing Variance and Bias of Treatment Effect Estimates: A Simulation Study,” American Journal of Epidemiology 190, no. 8 (2021): 1659–1670. S. Xu, C. Ross, M. A. Raebel, S. Shetterly, C. Blanchette, and D. Smith, “Use of Stabilized Inverse Propensity Scores as Weights to Directly Estimate Relative Risk and Its Confidence Intervals,” Value in Health 13, no. 2 (2010): 273–277. P. C. Austin and E. A. Stuart, “Moving Towards Best Practice When Using Inverse Probability of Treatment Weighting (IPTW) Using the Propensity Score to Estimate Causal Treatment Effects in Observational Studies,” Statistics in Medicine 34, no. 28 (2015): 3661–3679. K. Imai and M. Ratkovic, “Covariate Balancing Propensity Score,” Journal of the Royal Statistical Society. Series B, Statistical Methodology 76, no. 1 (2014): 243–263. Q. Zhao, “Covariate Balancing Propensity Score by Tailored Loss Functions,” Annals of Statistics 47, no. 2 (2019): 965–993. J. D. Huling and S. Mak, “Energy Balancing of Covariate Distributions,” Journal of Causal Inference 12, no. 1 (2024): 20220029. N. Kallus, “Generalized Optimal Matching Methods for Causal Inference,” Journal of Machine Learning Research 21, no. 62 (2020): 1–54. N. Kallus and M. Santacatterina, “Optimal Estimation of Generalized Average Treatment Effects Using Kernel Optimal Matching,” arXiv Preprint arXiv:190804748, (2019). J. M. Franklin, J. A. Rassen, D. Ackermann, D. B. Bartels, and S. Schneeweiss, “Metrics for Covariate Balance in Cohort Studies of Causal Effects,” Statistics in Medicine 33, no. 10 (2013): 1685–1699. J. M. Robins, A. Rotnitzky, and L. P. Zhao, “Estimation of Regression Coefficients When Some Regressors Are Not Always Observed,” Journal of the American Statistical Association 89, no. 427 (1994): 846–866. D. O. Scharfstein, A. Rotnitzky, and J. M. Robins, “Adjusting for Nonignorable Drop‐Out Using Semiparametric Nonresponse Models,” Journal of the American Statistical Association 94, no. 448 (1999): 1096–1120. A. Tsiatis, Semiparametric Theory and Missing Data, vol. 73 (Springer New York, 2006). M. J. van der Laan and D. Rubin, “Targeted Maximum Likelihood Learning,” International Journal of Biostatistics 2, no. 1 (2006): 1–40. D. B. Rubin, “Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies,” Journal of Educational Psychology 66, no. 5 (1974): 688–701. J. Neyman, “On the Application of Probability Theory to Agricultural Experiments. Essay on Principles. Section 9 (Translation Published in 1990),” Statistical Science 5 (1923): 472–480. J. Hájek, “Comment on “an Essay on the Logical Foundations of Survey Sampling, Part One” by D. Basu,” in Foundations of Statistical Inference, ed. V. P. Godambe and D. A. Sprott (Holt, Rinehart and Winston, 1971), 236–248. W. K. Newey and D. McFadden, “Chapter 36 Large Sample Estimation and Hypothesis Testing,” in Handbook of Econometrics, vol. 4 (Elsevier, 1994), 2111–2245. A. W. van der Vaart, Asymptotic Statistics. Cambridge Series in Statistical and Probabilistic Mathematics (Cambridge University Press, 1998). L. A. Stefanski and D. D. Boos, “The Calculus of M‐Estimation,” American Statistician 56, no. 1 (2002): 29–38. L. P. Hansen, “Large Sample Properties of Generalized Method of Moments Estimators,” Econometrica 50, no. 4 (1982): 1029–1054. C. Fong, M. Ratkovic, and K. Imai, “CBPS: Covariate Balancing Propensity Score,” Comprehensive R Archive Network, R package version 0.24, (2025), https://CRAN.R‐project.org/package=CBPS. G. J. Székely and M. L. Rizzo, “Energy Statistics: A Class of Statistics Based on Distances,” Journal of Statistical Planning and Inference 143, no. 8 (2013): 1249–1272. D. A. Freedman, “On the So‐Called “Huber Sandwich Estimator” and “Robust Standard Errors”,” American Statistician 60, no. 4 (2006): 299–302. J. D. Y. Kang and J. L. Schafer, “Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean From Incomplete Data,” Statistical Science 22, no. 4 (2007): 523–539. E. Peyrot, “Choosing Covariate Balancing Methods for Causal Inference: Practical Insights from a Simulation Study ‐ Analysis Code and Materials,” GitHub, (2025), https://github.com/EtiennePeyrot/benchmark_balancing_methods. N. Greifer, “WeightIt: Weighting for Covariate Balance in Observational Studies,” R package version 0.13.1. Comprehensive R Archive Network, (2022), https://CRAN.R‐project.org/package=WeightIt. Q. Zhao, Covalign: R Source Code for Covariate Balancing by Tailored Loss Functions (CBSR/TLF) (University of Cambridge, 2019), https://www.statslab.cam.ac.uk/qz280/publication/balancing‐loss/. A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola, “A Kernel Two‐Sample Test,” Journal of Machine Learning Research 13, no. 25 (2012): 723–773. D. Garreau, W. Jitkrittum, and M. Kanagawa, “Large Sample Analysis of the Median Heuristic,” arXiv Preprint arXiv:170707269, (2017). A. Ramdas, S. J. Reddi, B. Póczos, A. Singh, and L. Wasserman, “On the Decreasing Power of Kernel and Distance Based Nonparametric Hypothesis Tests in High Dimensions,” Proceedings of the AAAI Conference on Artificial Intelligence 29, no. 1 (2015): 3571–3577. A. Schrab, I. Kim, M. Albert, B. Laurent, B. Guedj, and A. Gretton, “MMD Aggregated Two‐Sample Test,” Journal of Machine Learning Research 24, no. 194 (2023): 1–81. N. Kallus, B. Pennicooke, and M. Santacatterina, “More Robust Estimation of Sample Average Treatment Effects Using Kernel Optimal Matching in an Observational Study of Spine Surgical Interventions,” arXiv Preprint arXiv:181104274, (2018), https://github.com/CausalML/KOM‐SATE. Gurobi Optimization, LLC, “Gurobi Optimizer Reference Manual,” Gurobi Optimization, LLC, (2024), https://www.gurobi.com. B. Stellato, G. Banjac, P. Goulart, and S. Boyd, “OSQP: Quadratic Programming Solver Using the OSQP Library,” R package version 0.6.0.5. Comprehensive R Archive Network; (2021), https://CRAN.R‐project.org/package=osqp. I. R. White, “Simsum: Analyses of Simulation Studies Including Monte Carlo Error,” Stata Journal 10, no. 3 (2010): 369–385. E. Koehler, E. Brown, and S. J. P. A. Haneuse, “On the Assessment of Monte Carlo Error in Simulation‐Based Statistical Analyses,” American Statistician 63, no. 2 (2009): 155–162. R Core Team, R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing, 2022), https://www.R‐project.org/. E. Goetghebeur, S. le Cessie, B. De Stavola, M. EEM, and I. Waernbaum, “Formulating Causal Questions and Principled Statistical Answers,” Statistics in Medicine 39, no. 30 (2020): 4922–4948. D. Sejdinovic, B. Sriperumbudur, A. Gretton, and K. Fukumizu, “Equivalence of Distance‐Based and RKHS‐Based Statistics in Hypothesis Testing,” Annals of Statistics 41, no. 5 (2013): 2263–2291. C. Leyrat, S. R. Seaman, I. R. White, et al., “Propensity Score Analysis With Partially Observed Covariates: How Should Multiple Imputation Be Used?,” Statistical Methods in Medical Research 28, no. 1 (2019): 3–19. J. M. Robins, M. A. Hernán, and B. Brumback, “Marginal Structural Models and Causal Inference in Epidemiology,” Epidemiology 11, no. 5 (2000): 550–560. G. W. Imbens, “The Role of the Propensity Score in Estimating Dose‐Response Functions,” Biometrika 87, no. 3 (2000): 706–710. K. Hirano and G. W. Imbens, “The Propensity Score With Continuous Treatments,” in Applied Bayesian Modeling and Causal Inference From Incomplete‐Data Perspectives, ed. A. Gelman and X. L. Meng (John Wiley & Sons, Ltd., 2004), 73–84. C. Fong, C. Hazlett, and K. Imai, “Covariate Balancing Propensity Score for a Continuous Treatment: Application to the Efficacy of Political Advertisements,” Annals of Applied Statistics 12, no. 1 (2018): 156–177. N. Kallus and M. Santacatterina, “Kernel Optimal Orthogonality Weighting: A Balancing Approach to Estimating Effects of Continuous Treatments,” arXiv Preprint arXiv:191011972, (2019). R. K. Crump, V. J. Hotz, G. W. Imbens, and O. A. Mitnik, “Dealing With Limited Overlap in Estimation of Average Treatment Effects,” Biometrika 96, no. 1 (2009): 187–199. B. K. Lee, J. Lessler, and E. A. Stuart, “Weight Trimming and Propensity Score Weighting,” PLoS One 6, no. 3 (2011): e18174. F. Li and L. E. Thomas, “Addressing Extreme Propensity Scores via the Overlap Weights,” American Journal of Epidemiology 188, no. 1 (2018): 250–257. |
| Grant Information: | ANR-22-CPJ1-0047-01 Agence Nationale de la Recherche; ANR-23-IACL-0008 Agence Nationale de la Recherche; ANR-18-CE36-0010-01 Agence Nationale de la Recherche |
| Contributed Indexing: | Keywords: Monte Carlo simulation; causal inference; inverse probability of treatment weighting; observational study; treatment effect estimation |
| Entry Date(s): | Date Created: 20260709 Date Completed: 20260709 Latest Revision: 20260726 |
| Update Code: | 20260726 |
| PubMed Central ID: | PMC13346537 |
| DOI: | 10.1002/sim.70672 |
| PMID: | 42420803 |
| Database: | MEDLINE |
| ISSN: | 1097-0258 |
|---|---|
| DOI: | 10.1002/sim.70672 |