| Literature DB >> 28137292 |
Mohammad W Hattab1, Andrey A Shabalin1, Shaunna L Clark1, Min Zhao1, Gaurav Kumar1, Robin F Chan1, Lin Ying Xie1, Rick Jansen2, Laura K M Han2, Patrik K E Magnusson3, Gerard van Grootheest2, Christina M Hultman3, Brenda W J H Penninx2, Karolina A Aberg1, Edwin J C G van den Oord4.
Abstract
Based on an extensive simulation study, McGregor and colleagues recently recommended the use of surrogate variable analysis (SVA) to control for the confounding effects of cell-type heterogeneity in DNA methylation association studies in scenarios where no cell-type proportions are available. As their recommendation was mainly based on simulated data, we sought to replicate findings in two large-scale empirical studies. In our empirical data, SVA did not fully correct for cell-type effects, its performance was somewhat unstable, and it carried a risk of missing true signals caused by removing variation that might be linked to actual disease processes. By contrast, a reference-based correction method performed well and did not show these limitations. A disadvantage of this approach is that if reference methylomes are not (publicly) available, they will need to be generated once for a small set of samples. However, given the notable risk we observed for cell-type confounding, we argue that, to avoid introducing false-positive findings into the literature, it could be well worth making this investment.Please see related Correspondence article: https://genomebiology.biomedcentral.com/articles/10/1186/s13059-017-1149-7 and related Research article: https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-0935-y.Entities:
Mesh:
Year: 2017 PMID: 28137292 PMCID: PMC5282865 DOI: 10.1186/s13059-017-1148-8
Source DB: PubMed Journal: Genome Biol ISSN: 1474-7596 Impact factor: 13.583
Comparison of reference-based and latent variable cell-type corrections in two empirical DNA methylation studies
| Depression MWAS study | Schizophrenia MWAS study | |||||||
|---|---|---|---|---|---|---|---|---|
| Enrich. | Enrich. | Number of | Increase | Enrich. | Enrich. | Number of | Increase | |
| No cell-type correction | 6.04 | <0.001 | – | – | 6.13 | <0.001 | – | – |
| Reference-based correction | 1.08 | 0.084 | – | 0.0% | 1.02 | 0.029 | – | 0.0% |
| SVA subset 1 | 1.11 | 0.001 | 84 | 8.1% | 6.54 | 0.001 | 5 | 0.8% |
| SVA subset 2 | 1.26 | 0.001 | 83 | 8.7% | 7.24 | <0.001 | 10 | 3.3% |
| SVA subset 3 | 1.85 | 0.004 | 19 | 2.0% | 6.45 | 0.001 | 5 | 0.9% |
| SVA subset 4 | 1.28 | <0.001 | 83 | 9.3% | 7.01 | 0.001 | 12 | 2.8% |
| SVA subset 5 | 3.05 | 0.001 | 14 | 1.9% | 6.79 | 0.002 | 6 | 1.2% |
| SVA subset 6 | 1.30 | 0.001 | 81 | 7.8% | 6.59 | <0.001 | 6 | 0.9% |
| SVA subset 7 | 1.07 | <0.001 | 78 | 9.2% | 6.42 | <0.001 | 4 | 0.7% |
| SVA subset 8 | 1.28 | 0.004 | 79 | 9.2% | 6.48 | <0.001 | 4 | 0.7% |
| SVA subset 9 | 1.13 | 0.003 | 84 | 9.2% | 7.46 | 0.001 | 10 | 2.7% |
| SVA subset 10 | 1.07 | 0.003 | 80 | 8.4% | 6.71 | 0.003 | 8 | 1.7% |
| SVA subset 11 | 1.06 | 0.006 | 21 | 2.7% | 6.48 | <0.001 | 8 | 1.3% |
| SVA subset 12 | 1.17 | 0.004 | 84 | 7.4% | 6.49 | 0.001 | 5 | 0.8% |
We used a permutation test to examine whether our top methylome-wide association study (MWAS) results were enriched for sites showing significant methylation differences among cell types. Our test preserved the correlation structure of the data by shifting the CpG coordinates of the case-control and cell-type MWAS by a single random number in each permutation. We examined multiple cut-offs (1, 5, and 10%) to define “top results” in the case-control and cell-type data and selected the most significant combination. We accounted for this “multiple testing” by also selecting the most significant finding in each permutation. The “No cell-type correction” model includes laboratory technical covariates, age, and sex. The other models include these same covariates where the “Reference-based correction” model adds estimates of cell-type proportions, and the SVA models add latent variables. “Enrich. ratio” is the ratio of the number of CpGs showing methylation differences between cell types among the top MWAS finding relative to the number expected under the null hypotheses assuming no enrichment; “Enrich. P value” is the probability under this null hypothesis, as determined through permutations; “Number of SVs” is the number of latent variables selected by SVA; “Increase r 2” is additional variance explained by SVs in case-control status compared with a multiple regression model that included technical covariates, age/sex, and cell-type estimates. SV surrogate variable, SVA surrogate variable analysis