| Literature DB >> 22373042 |
Abstract
Advances in next-generation sequencing technology are enabling researchers to capture a comprehensive picture of genomic variation across large numbers of individuals with unprecedented levels of efficiency. The main analytic challenge in disease mapping is how to mine the data for rare causal variants among a sea of neutral variation. To achieve this goal, investigators have proposed a number of methods that exploit biological knowledge. In this paper, I propose applying a Bayesian stochastic search variable selection algorithm in this context. My multivariate method is inspired by the combined multivariate and collapsing method. In this proposed method, however, I allow an arbitrary number of different sources of biological knowledge to inform the model as prior distributions in a two-level hierarchical model. This allows rare variants with similar prior distributions to share evidence of association. Using the 1000 Genomes Project single-nucleotide polymorphism data provided by Genetic Analysis Workshop 17, I show that through biologically informative prior distributions, some power can be gained over noninformative prior distributions.Entities:
Year: 2011 PMID: 22373042 PMCID: PMC3287850 DOI: 10.1186/1753-6561-5-S9-S16
Source DB: PubMed Journal: BMC Proc ISSN: 1753-6561
Posterior estimates of hyperparameters
| Parameter | Trait Q1 | Trait Q1 |
|---|---|---|
| 0.006 | 0.006 | |
| 0.01 | 0.01 | |
| 0.03 (0.06) | 0.02 (0.06) |
Accuracy of estimates of β for trait Q1 between maximum-likelihood estimate (MLE) and hierarchical modeling (HM) estimates
| Variable | Mean square error | Mean coverage ratea | ||
|---|---|---|---|---|
| MLE | HM | MLE | HM | |
| C1S6521 | 0.196 | 0.046 | 0.68 | 0.94 |
| C13S398 | 0.069 | 0.036 | 0.90 | 0.93 |
| C13S515 | 0.300 | 0.029 | 0.07 | 0.92 |
| C13S522 | 0.126 | 0.037 | 0.06 | 0.71 |
| HFE, nonsynonymous | 0.218 | 0.033 | 0.63 | 0.98 |
| KCTD14, nonsynonymous | 0.007 | 0.009 | 0.91 | 0.84 |
| C4S1878 | 0.102 | 0.024 | 0.7 | 0.95 |
a Proportion of replicates where true β falls within the 95% confidence interval.
Figure 1Receiver operating characteristic curve under polygenic disease model for trait Q1. The proportion of causal variants is plotted as a function of the proportion of noncausal variants, taken across 200 replicates.
Figure 2Receiver operating characteristic curve under polygenic disease model for traitQ2. The proportion of causal variants is plotted as a function of the proportion of noncausal variants, taken across 200 replicates.
Relative power (in relation to the maximum-likelihood estimate) of hierarchical modeling method at FDR = 0.05
| Variation | Trait Q1 | Trait Q2 |
|---|---|---|
| UNINF | 0.94 | 1.14 |
| 0.98 | 1.17 | |
| 1.04 | 1.17 | |
| FULL | 1.05 | 1.19 |
FDR, false discovery range.
Bayes factors for each causal variable under the Q1 trait model
| Causal variablea | UNINF | FULL | ||
|---|---|---|---|---|
| 1.14 | 1.95 | 1.85 | 3.99 | |
| C4S1884 | 36.33 | 43.31 | 36.12 | 47.65 |
| 0.58 | 1.14 | 0.57 | 1.22 | |
| C13S522 | 527.13 | 600.4 | 764.92 | 773.35 |
| C1S6533 | 109.38 | 149.45 | 119.39 | 162.42 |
| C4S1878 | 57.23 | 92.15 | 69.16 | 107.15 |
| C14S1734 | 1.47 | 2.42 | 1.47 | 2.58 |
| C13S431 | 299.20 | 327.49 | 572.87 | 551.06 |
| 17.25 | 27.80 | 81 | 103.17 | |
| C13S523 | 998.33 | 999.07 | 999.87 | 999.7 |
a Defined as either a bin of SNPs (shown with convention gene name and mutation class) or a single SNP. Only variables with MAF ≥ 0.01 were included for analyses.
Bayes factors for each causal variable under the Q2 trait model
| Causal variablea | UNINF | FULL | ||
|---|---|---|---|---|
| C6S5441 | 53.96 | 77.22 | 65.32 | 88.21 |
| 22.36 | 34.60 | 23.92 | 34.82 | |
| C2S354 | 11.17 | 15.29 | 11.04 | 16.46 |
| C8S442 | 62.14 | 88.72 | 64.60 | 87.26 |
| C6S5449 | 64.43 | 85.17 | 73.89 | 94.77 |
| 60.71 | 83.47 | 60.57 | 85.76 | |
| 49.64 | 71.03 | 49.12 | 70.54 | |
| C6S5426 | 0.87 | 1.48 | 1.25 | 2.10 |
| C6S5380 | 212.1 | 264.3 | 210.2 | 263.0 |
| 9.28 | 14.54 | 8.89 | 14.44 | |
| 17.51 | 26.85 | 17.15 | 26.93 | |
| 38.11 | 55.60 | 37.50 | 55.62 | |
| 1.34 | 2.43 | 1.62 | 2.63 |
a Defined as either a bin of SNPs (shown with convention gene name and mutation class) or a single SNP. Only variables with MAF ≥ 0.01 were included for analyses.