| Literature DB >> 19586935 |
Lin Wan1, Kelian Sun, Qi Ding, Yuehua Cui, Ming Li, Yalu Wen, Robert C Elston, Minping Qian, Wenjiang J Fu.
Abstract
Affymetrix SNP arrays have been widely used for single-nucleotide polymorphism (SNP) genotype calling and DNA copy number variation inference. Although numerous methods have <span class="Chemical">acpan>hieved high <span class="Chemical">accuracy in these fields, most studies have paid little attention to the modeling of hybridization of probes to off-target allele sequences, which can affect the accuracy greatly. In this study, we address this issue and demonstrate that hybridization with mismatch nucleotides (HWMMN) occurs in all SNP probe-sets and has a critical effect on the estimation of allelic concentrations (ACs). We study sequence binding through binding free energy and then binding affinity, and develop a probe intensity composite representation (PICR) model. The PICR model allows the estimation of ACs at a given SNP through statistical regression. Furthermore, we demonstrate with cell-line data of known true copy numbers that the PICR model can achieve reasonable accuracy in copy number estimation at a single SNP locus, by using the ratio of the estimated AC of each sample to that of the reference sample, and can reveal subtle genotype structure of SNPs at abnormal loci. We also demonstrate with HapMap data that the PICR model yields accurate SNP genotype calls consistently across samples, laboratories and even across array platforms.Entities:
Mesh:
Substances:
Year: 2009 PMID: 19586935 PMCID: PMC2761258 DOI: 10.1093/nar/gkp559
Source DB: PubMed Journal: Nucleic Acids Res ISSN: 0305-1048 Impact factor: 16.971
Figure 1.Comparison between raw intensities and predicted intensities by PICR of the 40 probes for three randomly selected SNPs of different genotypes on one randomly selected non-training Xba sample (NA10861_Xba_D12_4000189). (Top) SNP_A−1665980 of ‘AB’ type. (Middle) SNP_A−1756140 of ‘AA’ type. (Bottom) SNP_A−1652710 of ‘BB’ type.
Figure 2.Scatter plot of the allelic concentration estimated by PICR of all SNPs on one randomly selected non-training sample (NA07056_Xba_A11_4000090) with colors indicating HapMap genotype annotation.
Percentiles of the Pearson correlation (r) between the total concentration and true CNs
| Mean | 5% | 10% | 25% | 50% | 75% | 90% | 95% |
|---|---|---|---|---|---|---|---|
| 0.9688 | 0.8989 | 0.9351 | 0.9705 | 0.9868 | 0.9940 | 0.9971 | 0.9983 |
aThe total concentrations were computed at each of all 2361 SNPs on the X-chromosome of samples 1X, 2X, 3X, 4X and 5X (both Xba and Hind arrays).
Figure 3.Scatter plot of the total concentration estimated by PICR of all SNPs on the X-chromosome. (Upper left) Sample 1X versus sample 2X. (Upper right) Sample 3X versus sample 2X. (Lower left) Sample 4X versus sample 2X. (Lower right) Sample 5X versus sample 2X. The solid line in each panel represents the diagonal true concentration of sample 2X, and the dotted line represents the theoretical concentration of the sample (1X, 3X, 4X or 5X) relative to sample 2X. The closeness of the estimated concentration to the theoretical concentration illustrates the validity of this copy number estimation method by PICR. A slight bias to underestimation of the concentration was observed in the large copy number samples (4X, 5X).
Square root of Mean Squared Error (RtMSE) of the relative concentration estimated by PICR and mean intensity methods by sample of X-chromosomes
| Sample | 1X | 3X | 4X | 5X |
|---|---|---|---|---|
| True CN ratio | 0.5 | 1.5 | 2 | 2.5 |
| RtMSE By PICR | 0.1053 | 0.1766 | 0.2738 | 0.4730 |
| RtMSE by mean intensity | 0.2250 | 0.2613 | 0.5650 | 0.8467 |
Comparison between PICR and CRLMM in genotype-calling accuracy against the gold-standard HapMap genotype
| Sample | Genotype-calling method | Mean | SD | Median | 5% | 95% |
|---|---|---|---|---|---|---|
| 90 Xba arrays | PICR (single array) | 0.9966 | 0.0034 | 0.9975 | 0.9920 | 0.9987 |
| CRLMM (45 pairwise) | 0.9923 | 0.0021 | 0.9924 | 0.9886 | 0.9954 | |
| 15 Nsp arrays | PICR (single array) | 0.9916 | 0.0014 | 0.9921 | 0.9894 | 0.9930 |
| CRLMM (six pairs + one triplet) | 0.9806 | 0.0032 | 0.9813 | 0.9747 | 0.9844 | |
| CRLMM (all arrays together) | 0.9962 | 0.0003 | 0.9963 | 0.9958 | 0.9967 |
aStandard deviation.
b5th percentile.
c95th percentile.
Figure 4.The ROC curves for genotype calling on heterozygous SNPs versus homozygous SNPs by the PICR, the mean intensity and the CRLMM methods. The positive set contains all heterozygous SNPs by the HapMap annotation on the 90 HapMap Xba samples, while the negative set contains all homozygous SNPs on the same arrays. The ROC curves were obtained by varying the slopes c and 1/c of the lines used for clustering the SNPs into heterozygous or homozygous SNPs in the PICR and the mean intensity method, and by varying the cutoff value of the distance ratio ρ from the given SNP to the genotype centers in the CRLMM method (see ‘Materials and Methods’ section). (A) Comparison of ROC curve between the PICR and the mean intensity method. (B). Comparison of ROC curve between the PICR and the CRLMM method.
Figure 5.Comparison between the PICR method and the CRLMM method in genotype calling error rate using the HapMap genotype as the gold standard on 90 Xba arrays of the HapMap sample. The genotype calling was conducted with a pair of arrays each time by the CRLMM method, and with a single array by the PICR method.
Genotyping accuracy by CRLMM on HapMap samples in simultaneously calling 8 Xba arrays with varying number of HapMap samples
| HapMap + Study samples | 1+7 | 2+6 | 3+5 | 4+4 | 5+3 | 6+2 | 7+1 |
|---|---|---|---|---|---|---|---|
| Mean accuracy (%) On HapMap samples | 99.0 | 99.2 | 99.5 | 99.6 | 99.7 | 99.7 | 99.8 |
Figure 6.Estimated relative ACs of all 1204 SNPs located on the X-chromosome on the Xba array and their clusters. The ACs are relative to the corresponding SNP total concentration in the reference sample 2X. ‘NA’ is the AC of allele A, and ‘NB’ the AC of allele B. The colored dots represent the theoretical centers of the clusters of possible genotypes in each sample. (Left) Mean intensity method. (Right) the PICR method. The clustering was conducted by assigning each SNP to the genotype with the closest theoretical center. For example, in sample 4X relative to sample 2X, the relative true ACs at the theoretical centers are (2,0), (1.5, 0.5), (1,1), (0.5, 1.5) and (0,2), corresponding to the genotypes ‘AAAA’, ‘AAAB’, ‘AABB’, ‘ABBB’ and ‘BBBB’, respectively.