| Literature DB >> 22373088 |
Christopher E Schlosberg1, Tae-Hwi Schwantes-An, Weimin Duan, Nancy L Saccone.
Abstract
Using single-nucleotide polymorphism (SNP) genotypes from the 1000 Genomes Project pilot3 data provided for Genetic Analysis Workshop 17 (GAW17), we applied Bayesian network structure learning (BNSL) to identify potential causal SNPs associated with the Affected phenotype. We focus on the setting in which target genes that harbor causal variants have already been chosen for resequencing; the goal was to detect true causal SNPs from among the measured variants in these genes. Examining all available SNPs in the known causal genes, BNSL produced a Bayesian network from which subsets of SNPs connected to the Affected outcome were identified and measured for statistical significance using the hypergeometric distribution. The exploratory phase of analysis for pooled replicates sometimes identified a set of involved SNPs that contained more true causal SNPs than expected by chance in the Asian population. Analyses of single replicates gave inconsistent results. No nominally significant results were found in analyses of African or European populations. Overall, the method was not able to identify sets of involved SNPs that included a higher proportion of true causal SNPs than expected by chance alone. We conclude that this method, as currently applied, is not effective for identifying causal SNPs that follow the simulation model for the GAW17 data set, which includes many rare causal SNPs.Entities:
Year: 2011 PMID: 22373088 PMCID: PMC3287832 DOI: 10.1186/1753-6561-5-S9-S109
Source DB: PubMed Journal: BMC Proc ISSN: 1753-6561
Polymorphic SNPs genotyped in the three analyzed sample subsets
| Ancestry | Number of all polymorphic SNPs | Number of all polymorphic causal SNPs |
|---|---|---|
| African | 260 | 59 |
| Asian | 280 | 87 |
| European | 194 | 56 |
Figure 1Subset paradigms applied to identify candidate causal SNPs. Green represents causal SNPs, and red represents noncausal SNPs. The colored portions of the three right-hand panels show the descendants of Affected (DA), the Markov blanket of Affected (MBA), and the children of Affected (CA), respectively.
Primary results performed with log-likelihood scoring function on hill-climbing algorithm with 1,000 random restarts and 2,400 directional perturbations per score
| Ancestry | Replicate | Descendants of Affected (DA) | Markov blanket of Affected (MBA) | Children of Affected (DA ∩ MBA) | |||
|---|---|---|---|---|---|---|---|
| Number of causal SNPs/number of SNPs | Number of causal SNPs/number of SNPs | Number of causal SNPs/number of SNPs | |||||
| Asian | 1–10 | 12/25 | 11/45 | 0.891 | 2/3 | 0.228 | |
| 11–20 | 2/7 | 0.695 | 9/36 | 0.850 | 0/2 | 1 | |
| 1–20 | 69/182 | 15/55 | 0.798 | 2/7 | 0.695 | ||
| 21–40 | 15/69 | 0.983 | 24/63 | 0.113 | 2/4 | 0.367 | |
| 1 | 77/237 | 0.152 | 63/221 | 0.973 | 16/54 | 0.658 | |
| 2 | 75/235 | 0.305 | 55/221 | 0.999 | 17/64 | 0.851 | |
| 3 | 74/217 | 66/190 | 22/43 | ||||
| 4 | 76/227 | 55/190 | 0.894 | 11/42 | 0.821 | ||
| 5 | 76/240 | 0.371 | 66/218 | 0.758 | 25/68 | 0.155 | |
| 6 | 74/224 | 0.102 | 55/191 | 0.909 | 13/45 | 0.694 | |
| 7 | 67/203 | 0.160 | 60/202 | 0.826 | 16/55 | 0.693 | |
| 8 | 65/212 | 0.663 | 66/226 | 0.937 | 14/61 | 0.958 | |
| 9 | 73/230 | 0.368 | 67/226 | 0.887 | 24/86 | 0.816 | |
| 10 | 65/204 | 0.376 | 59/220 | 0.998 | 15/57 | 0.848 | |
| European | 1–10 | 34/105 | 0.155 | 11/46 | 0.849 | 1/6 | 0.874 |
| 11–20 | 6/11 | 6/22 | 0.655 | 2/2 | |||
| 1–20 | 0/1 | 1 | 2/16 | 0.972 | 0/1 | 0.711 | |
| 21–40 | 35/107 | 0.124 | 9/34 | 0.703 | 3/4 | ||
| African | 1–10 | 10/60 | 0.929 | 9/47 | 0.795 | 0/3 | 1 |
| 11–20 | 22/99 | 0.613 | 10/44 | 0.566 | 0/4 | 1 | |
| 1–20 | 7/34 | 0.695 | 11/53 | 0.707 | 0/4 | 1 | |
| 21–40 | 2/2 | 3/17 | 0.786 | 1/1 | 0.226 | ||
The analysis is for all typed, polymorphic SNPs in the target gene set. Probabilities less than 0.10 are shown in bold.