Combining Observational Studies to Reduce Multiple Biases.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Combining Observational Studies to Reduce Multiple Biases.
Συγγραφείς: Cole SR; From the Department of Epidemiology, UNC Gillings School of Global Public Health, Chapel Hill, NC., Zivich PN; From the Department of Epidemiology, UNC Gillings School of Global Public Health, Chapel Hill, NC., Shook-Sa BE; Department of Biostatistics, UNC Gillings School of Global Public Health, Chapel Hill, NC.; Nuffield Department of Population Health, University of Oxford, Oxford, United Kingdom., Richardson DB; Department of Environmental and Occupational Health, UC Irvine Wen School of Population and Public Health, Irvine, CA., Hudgens MG; Department of Biostatistics, UNC Gillings School of Global Public Health, Chapel Hill, NC., Edwards JK; From the Department of Epidemiology, UNC Gillings School of Global Public Health, Chapel Hill, NC.
Πηγή: Epidemiology (Cambridge, Mass.) [Epidemiology] 2026 Sep 01; Vol. 37 (5), pp. 565-574. Date of Electronic Publication: 2026 May 15.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Lippincott Williams & Wilkins Country of Publication: United States NLM ID: 9009644 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1531-5487 (Electronic) Linking ISSN: 10443983 NLM ISO Abbreviation: Epidemiology Subsets: MEDLINE
Imprint Name(s): Publication: <2000>- : Hagerstown, MD : Lippincott Williams & Wilkins
Original Publication: [Cambridge, MA : Blackwell Scientific Publications ; Chestnut Hill, MA : Epidemiology Resources, c1990-
Ιατρικοί όροι (MeSH): Observational Studies as Topic*/methods, Bias ; Monte Carlo Method ; Humans ; Confounding Factors, Epidemiologic
Περίληψη: Epidemiology stands to benefit greatly from combining data sources with complementary strengths. We propose a study design and estimators to combine information from multiple observational studies to simultaneously address confounding and outcome measurement error. Using inverse probability weighted, g-computation, and augmented inverse probability weighted estimators, we show how to combine information from two studies wherein the first study is subject to outcome misclassification (but has adequate confounder control) and the second study has gold-standard outcomes (but inadequate confounder control). Monte Carlo experiments demonstrate that the proposed estimators remove both confounding and measurement biases and provide appropriate 95% confidence interval coverage, while standard analyses fall short. Fusion designs offer a principled approach to combine data from multiple sources to address multiple biases in epidemiologic research.
(Copyright © 2026 Wolters Kluwer Health, LLC. All rights reserved.)
Competing Interests: Disclosure: The authors report no conflicts of interest.
References: Bareinboim E, Pearl J. Causal inference and the data-fusion problem. Proc Natl Acad Sci U S A. 2016;113:7345–7352.
Cole SR, Edwards JK, Breskin A, et al. Illustration of 2 fusion designs and estimators. Am J Epidemiol. 2023;192:467–474.
Li S, Luedtke A. Efficient estimation under data fusion. Biometrika. 2023;110:1041–1054.
Kezios KL, Zimmerman SC, Buto PT, et al. Overcoming data gaps in life course epidemiology by matching across cohorts. Epidemiology. 2024;35:610–617.
Graham E, Carone M, Rotnitzky A. Towards a unified theory for semiparametric data fusion with individual-level data. arXiv. 2024;2409:09973.
Kallus N, Mao X. On the role of surrogates in the efficient estimation of treatment effects with limited outcome data. J R Stat Soc Ser B Stat Methodol. 2025;87:480–509.
Li S, Gilbert PB, Duan R, Luedtke A. Data fusion using weakly aligned sources. J Am Stat Assoc. 2025;120:2569–2579.
Lawlor DA, Tilling K, Davey Smith G. Triangulation in aetiological epidemiology. Int J Epidemiol. 2016;45:1866–1886.
Greenland S, Morgenstern H. Confounding in health research. Annu Rev Public Health. 2001;22:189–212.
Spiegelman D. Approaches to uncertainty in exposure assessment in environmental epidemiology. Annu Rev Public Health. 2010;31:149–163.
Greenland S. Multiple-bias modelling for analysis of observational data. J R Stat Soc Ser A Stat Soc. 2005;168:267–306.
MacLehose RF, Ahern TP, Collin LJ, Li A, Lash TL. CYP2D6 phenotype and breast cancer outcomes: a bias analysis and meta-analysis. Cancer Epidemiol Biomarkers Prev. 2025;34:224–233.
Ahern TP, Collin LJ, MacLehose RF, et al. Adjusting adjustments: using external data to estimate the impact of different confounder sets on published associations. Epidemiology. 2025;36:381–390.
Post WS, Haberlen SA, Witt MD, et al. Suboptimal HIV suppression is associated with progression of coronary artery stenosis: the multicenter AIDS cohort study (MACS) longitudinal coronary CT angiography study. Atherosclerosis. 2022;353:33–40.
Crane HM, Paramsothy P, Drozd DR, et al.; Centers for AIDS Research Network of Integrated Clinical Systems (CNICS) Cohort. Types of myocardial infarction among human immunodeficiency virus-infected individuals in the United States. JAMA Cardiol. 2017;2:260–267.
Richardson DB, Rage E, Demers PA, et al. Mortality among uranium miners in North America and Europe: the pooled uranium miners analysis (PUMA). Int J Epidemiol. 2021;50:633–643.
Robins JM, Hernán MA, Brumback B. Marginal structural models and causal inference in epidemiology. Epidemiology. 2000;11:550–560.
Hernán MA, Cole SR. Invited commentary: causal diagrams and measurement bias. Am J Epidemiol. 2009;170:959–962.
Aronow PM, Robins JM, Saarinen T, Sävje F, Sekhon J. Nonparametric identification is not enough, but randomized controlled trials are. arXiv. 2021;2108:11342.
Rogan WJ, Gladen B. Estimating prevalence from the results of a screening test. Am J Epidemiol. 1978;107:71–76.
Stefanski LA, Boos DD. The calculus of M-estimation. Am Stat. 2002;56:29–38.
Levenberg K. A method for the solution of certain non-linear problems in least squares. Q Appl Math. 1944;2:164–168.
Cole SR, Breskin A, Shook-Sa BE, Zivich PN, Hudgens MG, Edwards JK. Five facts about influence functions. Epidemiology. 2025;36:467–472.
Cole SR, Edwards JK, Greenland S. Surprise! Am J Epidemiol. 2021;190:191–193.
Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Stat Med. 2019;38:2074–2102.
Ogburn EL, Rudolph KE, Morello-Frosch R, Khan A, Casey JA. A warning about using predicted values from regression models for epidemiologic inquiry. Am J Epidemiol. 2021;190:1142–1147.
Westreich D, Edwards JK, Lesko CR, Stuart E, Cole SR. Transportability of trial results using inverse odds of sampling weights. Am J Epidemiol. 2017;186:1010–1014.
Lesko CR, Buchanan AL, Westreich D, Edwards JK, Hudgens MG, Cole SR. Generalizing study results: a potential outcomes perspective. Epidemiology. 2017;28:553–561.
Godambe VP. Estimating Functions. Clarendon Press; 1991.
Naimi AI, Balzer LB. Stacked generalization: an introduction to super learning. Eur J Epidemiol. 2018;33:459–464.
Imbens GW. Nonparametric estimation of average treatment effects under exogeneity: a review. Rev Econ Stat. 2004;86:4–29.
Zivich PN, Breskin A. Machine learning for causal inference: on the use of cross-fit estimators. Epidemiology. 2021;32:393–401.
Alam S, Moodie EEM, Stephens DA. Should a propensity score model be super? The utility of ensemble procedures for causal adjustment. Stat Med. 2019;38:1690–1702.
Keele LJ, Small D. Comparing covariate prioritization via matching to machine learning methods for causal inference using five empirical applications. Am Stat. 2021;75:355–363.
Pirracchio R, Petersen ML, van der Laan M. Improving propensity score estimators’ robustness to model misspecification using super learner. Am J Epidemiol. 2015;181:108–119.
Dorie V, Hill J, Shalit U, Scott M, Cervone D. Automated versus do-it-yourself methods for causal inference: lessons learned from a data analysis competition. Stat Sci. 2019;34:43–68.
Rudolph KE, Williams NT, Miles CH, Antonelli J, Diaz I. All models are wrong, but which are useful? Comparing parametric and nonparametric estimation of causal effects in finite samples. J Causal Inference. 2023;11:20230022.
Godambe VP. Estimating functions: a synthesis of least squares and maximum likelihood. Inst Math Stat Lecture Notes Monogr Ser. 1997;32:5–16.
Cole SR, Edwards JK, Breskin A, et al. Respond to combining information from diverse sources. Am J Epidemiol. 2024;193:751–752.
Breskin A, Cole SR, Edwards JK, Brookmeyer R, Eron JJ, Adimora AA. Fusion designs and estimators for treatment effects. Stat Med. 2021;40:3124–3137.
Shook-Sa BE, Zivich PN, Rosin SP, et al. Fusing trial data for treatment comparisons: single vs multi-span bridging. Stat Med. 2024;43:793–815.
Cole SR, Shook-Sa BE, Zivich PN, Edwards JK, Richardson DB, Hudgens MG. Higher-order evidence. Eur J Epidemiol. 2024;39:1–11.
Grant Information: K01 AI177102 United States AI NIAID NIH HHS; K01 AI182506 United States AI NIAID NIH HHS; P30 AI050410 United States AI NIAID NIH HHS; R01 AI157758 United States AI NIAID NIH HHS
Contributed Indexing: Keywords: Bias; Cohort Study; Confounding; Data Fusion; Outcome Measurement Error; Random Error
Entry Date(s): Date Created: 20260515 Date Completed: 20260729 Latest Revision: 20260801
Update Code: 20260801
PubMed Central ID: PMC13182983
DOI: 10.1097/EDE.0000000000001999
PMID: 42138357
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:1531-5487
DOI:10.1097/EDE.0000000000001999