Literature DB >> 28137292

Correcting for cell-type effects in DNA methylation studies: reference-based method outperforms latent variable approaches in empirical studies.

Mohammad W Hattab1, Andrey A Shabalin1, Shaunna L Clark1, Min Zhao1, Gaurav Kumar1, Robin F Chan1, Lin Ying Xie1, Rick Jansen2, Laura K M Han2, Patrik K E Magnusson3, Gerard van Grootheest2, Christina M Hultman3, Brenda W J H Penninx2, Karolina A Aberg1, Edwin J C G van den Oord4.   

Abstract

Based on an extensive simulation study, McGregor and colleagues recently recommended the use of surrogate variable analysis (SVA) to control for the confounding effects of cell-type heterogeneity in DNA methylation association studies in scenarios where no cell-type proportions are available. As their recommendation was mainly based on simulated data, we sought to replicate findings in two large-scale empirical studies. In our empirical data, SVA did not fully correct for cell-type effects, its performance was somewhat unstable, and it carried a risk of missing true signals caused by removing variation that might be linked to actual disease processes. By contrast, a reference-based correction method performed well and did not show these limitations. A disadvantage of this approach is that if reference methylomes are not (publicly) available, they will need to be generated once for a small set of samples. However, given the notable risk we observed for cell-type confounding, we argue that, to avoid introducing false-positive findings into the literature, it could be well worth making this investment.Please see related Correspondence article: https://genomebiology.biomedcentral.com/articles/10/1186/s13059-017-1149-7 and related Research article: https://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-0935-y.

Entities:  

Mesh:

Year:  2017        PMID: 28137292      PMCID: PMC5282865          DOI: 10.1186/s13059-017-1148-8

Source DB:  PubMed          Journal:  Genome Biol        ISSN: 1474-7596            Impact factor:   13.583


Correspondence

Tissues often consist of multiple cell types that show different methylation patterns. In association studies, these differences can cause spurious findings when the relative abundance of the cell types is related to the outcome of interest. The inclusion of cell-type proportions as covariates will prevent such false positives. To avoid performing cell counts on all subjects in the study, these proportions can be estimated by using a small set of reference methylomes obtained using DNA from sorted cells [1]. However, reference methylomes might not always be (publicly) available and or be difficult to generate. In these scenarios, latent variables obtained by a decomposition of the methylation data can be used as a proxy for cell-type proportions. McGregor et al. [2] performed an extensive simulation study comparing one reference-based and seven latent variable methods. Although not always the best method, the reference-based method performed well. For scenarios where no reference is available, the authors recommended the use of surrogate variable analysis (SVA) [3], which performed adequately in all simulation scenarios. As the recommendation by McGregor and colleagues [2] was based mainly on simulated data, we studied SVA in two large-scale empirical studies. The first involved 1149 Dutch subjects (825 cases with depression and 324 controls) aged 18–65 years [4] and the second 1448 Swedish subjects (774 schizophrenia cases and 674 controls) aged 25–92 years [5, 6]. Using whole-blood samples from six US subjects, cell populations were isolated by positive selection using EasySep™ kits (Stemcell Technologies), which apply magnetic nanoparticles coated with antibodies against a particular surface antigen (CD molecules). Specifically, we used CD3, CD19, CD20, CD14, and CD15 to isolate all common cell types in blood. All methylation data were generated using methyl-CG binding domain sequencing (MBD-seq) [7, 8], but the schizophrenia study was conducted on an older sequencing platform with a slightly different laboratory protocol. We used a permutation test to examine whether our top methylome-wide association study (MWAS) results were enriched for sites showing significant methylation differences among cell types. The MBD-seq procedure assays almost all 28 million common CpGs in the human genome. As the SVA package could not process all sites simultaneously, it was performed on 12 randomly selected subsets of 100,000 CpG sites. Table 1 indicates that, if no cell-type correction is applied, MWAS findings show a greater than sixfold enrichment of CpG sites exhibiting cell-type differences in methylation. This was consistent with the significant case-control differences in estimated cell-type proportions (across cell types/studies, the median P value was 8.0 × 10–5) and stresses the need to control for this confounder. The enrichment disappears when using the reference-based method. By contrast, significant enrichment remained after SVA correction in all studied scenarios. The performance of SVA was associated with the number of surrogate variables (SVs), which varied considerably across the 12 randomly selected CpG subsets within each study. However, even when as many as 84 SVs were included, SVA failed to control for more-subtle cell-type effects. To enable a simultaneous analysis of all sites, analyses were repeated using principal component analysis (PCA) [9], which also corrects for cell types by using latent variables. However, this did not improve results.
Table 1

Comparison of reference-based and latent variable cell-type corrections in two empirical DNA methylation studies

Depression MWAS studySchizophrenia MWAS study
Enrich.ratioEnrich. P valueNumber ofSVsIncrease r 2 Enrich.ratioEnrich. P valueNumber ofSVsIncrease r 2
No cell-type correction6.04<0.0016.13<0.001
Reference-based correction1.080.0840.0%1.020.0290.0%
SVA subset 11.110.001848.1%6.540.00150.8%
SVA subset 21.260.001838.7%7.24<0.001103.3%
SVA subset 31.850.004192.0%6.450.00150.9%
SVA subset 41.28<0.001839.3%7.010.001122.8%
SVA subset 53.050.001141.9%6.790.00261.2%
SVA subset 61.300.001817.8%6.59<0.00160.9%
SVA subset 71.07<0.001789.2%6.42<0.00140.7%
SVA subset 81.280.004799.2%6.48<0.00140.7%
SVA subset 91.130.003849.2%7.460.001102.7%
SVA subset 101.070.003808.4%6.710.00381.7%
SVA subset 111.060.006212.7%6.48<0.00181.3%
SVA subset 121.170.004847.4%6.490.00150.8%

We used a permutation test to examine whether our top methylome-wide association study (MWAS) results were enriched for sites showing significant methylation differences among cell types. Our test preserved the correlation structure of the data by shifting the CpG coordinates of the case-control and cell-type MWAS by a single random number in each permutation. We examined multiple cut-offs (1, 5, and 10%) to define “top results” in the case-control and cell-type data and selected the most significant combination. We accounted for this “multiple testing” by also selecting the most significant finding in each permutation. The “No cell-type correction” model includes laboratory technical covariates, age, and sex. The other models include these same covariates where the “Reference-based correction” model adds estimates of cell-type proportions, and the SVA models add latent variables. “Enrich. ratio” is the ratio of the number of CpGs showing methylation differences between cell types among the top MWAS finding relative to the number expected under the null hypotheses assuming no enrichment; “Enrich. P value” is the probability under this null hypothesis, as determined through permutations; “Number of SVs” is the number of latent variables selected by SVA; “Increase r 2” is additional variance explained by SVs in case-control status compared with a multiple regression model that included technical covariates, age/sex, and cell-type estimates. SV surrogate variable, SVA surrogate variable analysis

Comparison of reference-based and latent variable cell-type corrections in two empirical DNA methylation studies We used a permutation test to examine whether our top methylome-wide association study (MWAS) results were enriched for sites showing significant methylation differences among cell types. Our test preserved the correlation structure of the data by shifting the CpG coordinates of the case-control and cell-type MWAS by a single random number in each permutation. We examined multiple cut-offs (1, 5, and 10%) to define “top results” in the case-control and cell-type data and selected the most significant combination. We accounted for this “multiple testing” by also selecting the most significant finding in each permutation. The “No cell-type correction” model includes laboratory technical covariates, age, and sex. The other models include these same covariates where the “Reference-based correction” model adds estimates of cell-type proportions, and the SVA models add latent variables. “Enrich. ratio” is the ratio of the number of CpGs showing methylation differences between cell types among the top MWAS finding relative to the number expected under the null hypotheses assuming no enrichment; “Enrich. P value” is the probability under this null hypothesis, as determined through permutations; “Number of SVs” is the number of latent variables selected by SVA; “Increase r 2” is additional variance explained by SVs in case-control status compared with a multiple regression model that included technical covariates, age/sex, and cell-type estimates. SV surrogate variable, SVA surrogate variable analysis The use of a reference-based method ensures that only variation linked to differences in cell-type proportions is eliminated. SVA can eliminate any general source of variation in the methylation data. This carries the risk of missing true signals when some SVs capture part of the disease processes (e.g., a pathway). Table 1 reports additional variance explained by SVs in case-control status compared with a multiple-regression model that included technical covariates, age/sex, and cell-type proportions. Depending on the number of SVs, the additional variance ranged from 1 to 9%. This illustrates the risk of SVA potentially eliminating true signals in a MWAS. To mitigate this risk, one could avoid regressing out SVs associated with the case-control status. However, as cell-type proportions are related to both case-control status and SVs, such a modified analysis might be even less effective in controlling for cell-type effects. With empirical data, SVA did not adequately correct for cell-type effects, had somewhat unstable performance, and carried a risk of missing true disease signals. The PCA suggested that these limitations might not be specific to SVA but are inherent to the use of latent variables—that is, whereas these corrections assume that cell-type heterogeneity impacts many sites, cell-type effects seem more subtle and cannot be fully captured by just the main latent variables. For this reason, we expect our findings to generalize to methylation platforms other than MBD-seq. By contrast, the reference-based method was superior in all respects. If reference methylomes are not (publicly) available for a given tissue and methylation assay, they will need to be generated once for a small set of samples. However, given the notable risk we observed for cell-type confounding, to avoid introducing false-positive findings into the literature it could be well worth making this investment.
  9 in total

1.  Methylome-wide association study of schizophrenia: identifying blood biomarker signatures of environmental insults.

Authors:  Karolina A Aberg; Joseph L McClay; Srilaxmi Nerella; Shaunna Clark; Gaurav Kumar; Wenan Chen; Amit N Khachane; Linying Xie; Alexandra Hudson; Guimin Gao; Aki Harada; Christina M Hultman; Patrick F Sullivan; Patrik K E Magnusson; Edwin J C G van den Oord
Journal:  JAMA Psychiatry       Date:  2014-03       Impact factor: 21.596

2.  MBD-seq as a cost-effective approach for methylome-wide association studies: demonstration in 1500 case--control samples.

Authors:  Karolina A Aberg; Joseph L McClay; Srilaxmi Nerella; Lin Y Xie; Shaunna L Clark; Alexandra D Hudson; Jozsef Bukszár; Daniel Adkins; Christina M Hultman; Patrick F Sullivan; Patrik K E Magnusson; Edwin J C G van den Oord
Journal:  Epigenomics       Date:  2012-12       Impact factor: 4.778

3.  The Netherlands Study of Depression and Anxiety (NESDA): rationale, objectives and methods.

Authors:  Brenda W J H Penninx; Aartjan T F Beekman; Johannes H Smit; Frans G Zitman; Willem A Nolen; Philip Spinhoven; Pim Cuijpers; Peter J De Jong; Harm W J Van Marwijk; Willem J J Assendelft; Klaas Van Der Meer; Peter Verhaak; Michel Wensing; Ron De Graaf; Witte J Hoogendijk; Johan Ormel; Richard Van Dyck
Journal:  Int J Methods Psychiatr Res       Date:  2008       Impact factor: 4.035

4.  Evaluation of Methyl-Binding Domain Based Enrichment Approaches Revisited.

Authors:  Karolina A Aberg; Linying Xie; Robin F Chan; Min Zhao; Ashutosh K Pandey; Gaurav Kumar; Shaunna L Clark; Edwin J C G van den Oord
Journal:  PLoS One       Date:  2015-07-15       Impact factor: 3.240

5.  MethylPCA: a toolkit to control for confounders in methylome-wide association studies.

Authors:  Wenan Chen; Guimin Gao; Srilaxmi Nerella; Christina M Hultman; Patrik K E Magnusson; Patrick F Sullivan; Karolina A Aberg; Edwin J C G van den Oord
Journal:  BMC Bioinformatics       Date:  2013-03-02       Impact factor: 3.169

6.  DNA methylation arrays as surrogate measures of cell mixture distribution.

Authors:  Eugene Andres Houseman; William P Accomando; Devin C Koestler; Brock C Christensen; Carmen J Marsit; Heather H Nelson; John K Wiencke; Karl T Kelsey
Journal:  BMC Bioinformatics       Date:  2012-05-08       Impact factor: 3.169

7.  Capturing heterogeneity in gene expression studies by surrogate variable analysis.

Authors:  Jeffrey T Leek; John D Storey
Journal:  PLoS Genet       Date:  2007-08-01       Impact factor: 5.917

8.  Genome-wide association analysis identifies 13 new risk loci for schizophrenia.

Authors:  Stephan Ripke; Colm O'Dushlaine; Kimberly Chambert; Jennifer L Moran; Anna K Kähler; Susanne Akterin; Sarah E Bergen; Ann L Collins; James J Crowley; Menachem Fromer; Yunjung Kim; Sang Hong Lee; Patrik K E Magnusson; Nick Sanchez; Eli A Stahl; Stephanie Williams; Naomi R Wray; Kai Xia; Francesco Bettella; Anders D Borglum; Brendan K Bulik-Sullivan; Paul Cormican; Nick Craddock; Christiaan de Leeuw; Naser Durmishi; Michael Gill; Vera Golimbet; Marian L Hamshere; Peter Holmans; David M Hougaard; Kenneth S Kendler; Kuang Lin; Derek W Morris; Ole Mors; Preben B Mortensen; Benjamin M Neale; Francis A O'Neill; Michael J Owen; Milica Pejovic Milovancevic; Danielle Posthuma; John Powell; Alexander L Richards; Brien P Riley; Douglas Ruderfer; Dan Rujescu; Engilbert Sigurdsson; Teimuraz Silagadze; August B Smit; Hreinn Stefansson; Stacy Steinberg; Jaana Suvisaari; Sarah Tosato; Matthijs Verhage; James T Walters; Douglas F Levinson; Pablo V Gejman; Kenneth S Kendler; Claudine Laurent; Bryan J Mowry; Michael C O'Donovan; Michael J Owen; Ann E Pulver; Brien P Riley; Sibylle G Schwab; Dieter B Wildenauer; Frank Dudbridge; Peter Holmans; Jianxin Shi; Margot Albus; Madeline Alexander; Dominique Campion; David Cohen; Dimitris Dikeos; Jubao Duan; Peter Eichhammer; Stephanie Godard; Mark Hansen; F Bernard Lerer; Kung-Yee Liang; Wolfgang Maier; Jacques Mallet; Deborah A Nertney; Gerald Nestadt; Nadine Norton; Francis A O'Neill; George N Papadimitriou; Robert Ribble; Alan R Sanders; Jeremy M Silverman; Dermot Walsh; Nigel M Williams; Brandon Wormley; Maria J Arranz; Steven Bakker; Stephan Bender; Elvira Bramon; David Collier; Benedicto Crespo-Facorro; Jeremy Hall; Conrad Iyegbe; Assen Jablensky; Rene S Kahn; Luba Kalaydjieva; Stephen Lawrie; Cathryn M Lewis; Kuang Lin; Don H Linszen; Ignacio Mata; Andrew McIntosh; Robin M Murray; Roel A Ophoff; John Powell; Dan Rujescu; Jim Van Os; Muriel Walshe; Matthias Weisbrod; Durk Wiersma; Peter Donnelly; Ines Barroso; Jenefer M Blackwell; Elvira Bramon; Matthew A Brown; Juan P Casas; Aiden P Corvin; Panos Deloukas; Audrey Duncanson; Janusz Jankowski; Hugh S Markus; Christopher G Mathew; Colin N A Palmer; Robert Plomin; Anna Rautanen; Stephen J Sawcer; Richard C Trembath; Ananth C Viswanathan; Nicholas W Wood; Chris C A Spencer; Gavin Band; Céline Bellenguez; Colin Freeman; Garrett Hellenthal; Eleni Giannoulatou; Matti Pirinen; Richard D Pearson; Amy Strange; Zhan Su; Damjan Vukcevic; Peter Donnelly; Cordelia Langford; Sarah E Hunt; Sarah Edkins; Rhian Gwilliam; Hannah Blackburn; Suzannah J Bumpstead; Serge Dronov; Matthew Gillman; Emma Gray; Naomi Hammond; Alagurevathi Jayakumar; Owen T McCann; Jennifer Liddle; Simon C Potter; Radhi Ravindrarajah; Michelle Ricketts; Avazeh Tashakkori-Ghanbaria; Matthew J Waller; Paul Weston; Sara Widaa; Pamela Whittaker; Ines Barroso; Panos Deloukas; Christopher G Mathew; Jenefer M Blackwell; Matthew A Brown; Aiden P Corvin; Mark I McCarthy; Chris C A Spencer; Elvira Bramon; Aiden P Corvin; Michael C O'Donovan; Kari Stefansson; Edward Scolnick; Shaun Purcell; Steven A McCarroll; Pamela Sklar; Christina M Hultman; Patrick F Sullivan
Journal:  Nat Genet       Date:  2013-08-25       Impact factor: 38.330

9.  An evaluation of methods correcting for cell-type heterogeneity in DNA methylation studies.

Authors:  Kevin McGregor; Sasha Bernatsky; Ines Colmegna; Marie Hudson; Tomi Pastinen; Aurélie Labbe; Celia M T Greenwood
Journal:  Genome Biol       Date:  2016-05-03       Impact factor: 13.583

  9 in total
  11 in total

1.  Epigenetic Aging in Major Depressive Disorder.

Authors:  Laura K M Han; Moji Aghajani; Shaunna L Clark; Robin F Chan; Mohammad W Hattab; Andrey A Shabalin; Min Zhao; Gaurav Kumar; Lin Ying Xie; Rick Jansen; Yuri Milaneschi; Brian Dean; Karolina A Aberg; Edwin J C G van den Oord; Brenda W J H Penninx
Journal:  Am J Psychiatry       Date:  2018-04-16       Impact factor: 18.112

Review 2.  Accelerating research on biological aging and mental health: Current challenges and future directions.

Authors:  Laura K M Han; Josine E Verhoeven; Audrey R Tyrka; Brenda W J H Penninx; Owen M Wolkowitz; Kristoffer N T Månsson; Daniel Lindqvist; Marco P Boks; Dóra Révész; Synthia H Mellon; Martin Picard
Journal:  Psychoneuroendocrinology       Date:  2019-04-05       Impact factor: 4.905

3.  Independent Methylome-Wide Association Studies of Schizophrenia Detect Consistent Case-Control Differences.

Authors:  Robin F Chan; Andrey A Shabalin; Carolina Montano; Eilis Hannon; Christina M Hultman; Margaret D Fallin; Andrew P Feinberg; Jonathan Mill; Edwin J C G van den Oord; Karolina A Aberg
Journal:  Schizophr Bull       Date:  2020-02-26       Impact factor: 9.306

Review 4.  Statistical and integrative system-level analysis of DNA methylation data.

Authors:  Andrew E Teschendorff; Caroline L Relton
Journal:  Nat Rev Genet       Date:  2017-11-13       Impact factor: 53.242

5.  In silico deconvolution and purification of cancer epigenomes.

Authors:  Martí Duran-Ferrer; Renée Beekman; José I Martín-Subero
Journal:  Oncoscience       Date:  2017-04-14

6.  Response to: Correcting for cell-type effects in DNA methylation studies: reference-based method outperforms latent variable approaches in empirical studies.

Authors:  Kevin McGregor; Aurélie Labbe; Celia M T Greenwood
Journal:  Genome Biol       Date:  2017-01-30       Impact factor: 13.583

Review 7.  Maximizing ecological and evolutionary insight in bisulfite sequencing data sets.

Authors:  Amanda J Lea; Tauras P Vilgalys; Paul A P Durst; Jenny Tung
Journal:  Nat Ecol Evol       Date:  2017-07-21       Impact factor: 15.460

8.  Convergence of evidence from a methylome-wide CpG-SNP association study and GWAS of major depressive disorder.

Authors:  Karolina A Aberg; Andrey A Shabalin; Robin F Chan; Min Zhao; Gaurav Kumar; Gerard van Grootheest; Shaunna L Clark; Lin Y Xie; Yuri Milaneschi; Brenda W J H Penninx; Edwin J C G van den Oord
Journal:  Transl Psychiatry       Date:  2018-08-22       Impact factor: 6.222

9.  Epigenetic aging is accelerated in alcohol use disorder and regulated by genetic variation in APOL2.

Authors:  Audrey Luo; Jeesun Jung; Martha Longley; Daniel B Rosoff; Katrin Charlet; Christine Muench; Jisoo Lee; Colin A Hodgkinson; David Goldman; Steve Horvath; Zachary A Kaminsky; Falk W Lohoff
Journal:  Neuropsychopharmacology       Date:  2019-08-29       Impact factor: 7.853

10.  Methylome-wide association findings for major depressive disorder overlap in blood and brain and replicate in independent brain samples.

Authors:  Karolina A Aberg; Brian Dean; Andrey A Shabalin; Robin F Chan; Laura K M Han; Min Zhao; Gerard van Grootheest; Lin Y Xie; Yuri Milaneschi; Shaunna L Clark; Gustavo Turecki; Brenda W J H Penninx; Edwin J C G van den Oord
Journal:  Mol Psychiatry       Date:  2018-09-21       Impact factor: 15.992

View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.