Literature DB >> 31346355

ESTIMATING POPULATION AVERAGE CAUSAL EFFECTS IN THE PRESENCE OF NON-OVERLAP: THE EFFECT OF NATURAL GAS COMPRESSOR STATION EXPOSURE ON CANCER MORTALITY.

Rachel C Nethery1, Fabrizia Mealli2, Francesca Dominici1.   

Abstract

Most causal inference studies rely on the assumption of overlap to estimate population or sample average causal effects. When data suffer from non-overlap, estimation of these estimands requires reliance on model specifications, due to poor data support. All existing methods to address non-overlap, such as trimming or down-weighting data in regions of poor data support, change the estimand so that inference cannot be made on the sample or the underlying population. In environmental health research settings, where study results are often intended to influence policy, population-level inference may be critical, and changes in the estimand can diminish the impact of the study results, because estimates may not be representative of effects in the population of interest to policymakers. Researchers may be willing to make additional, minimal modeling assumptions in order to preserve the ability to estimate population average causal effects. We seek to make two contributions on this topic. First, we propose a flexible, data-driven definition of propensity score overlap and non-overlap regions. Second, we develop a novel Bayesian framework to estimate population average causal effects with minor model dependence and appropriately large uncertainties in the presence of non-overlap and causal effect heterogeneity. In this approach, the tasks of estimating causal effects in the overlap and non-overlap regions are delegated to two distinct models, suited to the degree of data support in each region. Tree ensembles are used to non-parametrically estimate individual causal effects in the overlap region, where the data can speak for themselves. In the non-overlap region, where insufficient data support means reliance on model specification is necessary, individual causal effects are estimated by extrapolating trends from the overlap region via a spline model. The promising performance of our method is demonstrated in simulations. Finally, we utilize our method to perform a novel investigation of the causal effect of natural gas compressor station exposure on cancer outcomes. Code and data to implement the method and reproduce all simulations and analyses is available on Github (https://github.com/rachelnethery/overlap).

Entities:  

Keywords:  Bayesian Additive Regression Trees; Cancer Mortality; Natural Gas; Overlap; Propensity Score; Splines

Year:  2019        PMID: 31346355      PMCID: PMC6658123          DOI: 10.1214/18-AOAS1231

Source DB:  PubMed          Journal:  Ann Appl Stat        ISSN: 1932-6157            Impact factor:   2.083


  2 in total

1.  In Pursuit of Evidence in Air Pollution Epidemiology: The Role of Causally Driven Data Science.

Authors:  Marco Carone; Francesca Dominici; Lianne Sheppard
Journal:  Epidemiology       Date:  2020-01       Impact factor: 4.822

Review 2.  Core concepts in pharmacoepidemiology: Violations of the positivity assumption in the causal analysis of observational data: Consequences and statistical approaches.

Authors:  Yaqian Zhu; Rebecca A Hubbard; Jessica Chubak; Jason Roy; Nandita Mitra
Journal:  Pharmacoepidemiol Drug Saf       Date:  2021-08-24       Impact factor: 2.890

  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.