| Literature DB >> 21249187 |
Dandan Zhang1, Yudong Qian, Nirmala Akula, Ney Alliey-Rodriguez, Jinsong Tang, Elliot S Gershon, Chunyu Liu.
Abstract
Several computer programs are available for detecting copy number variants (CNVs) using genome-wide SNP arrays. We evaluated the performance of four CNV detection software suites--Birdsuite, Partek, HelixTree, and PennCNV-Affy--in the identification of both rare and common CNVs. Each program's performance was assessed in two ways. The first was its recovery rate, i.e., its ability to call 893 CNVs previously identified in eight HapMap samples by paired-end sequencing of whole-genome fosmid clones, and 51,440 CNVs identified by array Comparative Genome Hybridization (aCGH) followed by validation procedures, in 90 HapMap CEU samples. The second evaluation was program performance calling rare and common CNVs in the Bipolar Genome Study (BiGS) data set (1001 bipolar cases and 1033 controls, all of European ancestry) as measured by the Affymetrix SNP 6.0 array. Accuracy in calling rare CNVs was assessed by positive predictive value, based on the proportion of rare CNVs validated by quantitative real-time PCR (qPCR), while accuracy in calling common CNVs was assessed by false positive/false negative rates based on qPCR validation results from a subset of common CNVs. Birdsuite recovered the highest percentages of known HapMap CNVs containing >20 markers in two reference CNV datasets. The recovery rate increased with decreased CNV frequency. In the tested rare CNV data, Birdsuite and Partek had higher positive predictive values than the other software suites. In a test of three common CNVs in the BiGS dataset, Birdsuite's call was 98.8% consistent with qPCR quantification in one CNV region, but the other two regions showed an unacceptable degree of accuracy. We found relatively poor consistency between the two "gold standards," the sequence data of Kidd et al., and aCGH data of Conrad et al. Algorithms for calling CNVs especially common ones need substantial improvement, and a "gold standard" for detection of CNVs remains to be established.Entities:
Mesh:
Year: 2011 PMID: 21249187 PMCID: PMC3020939 DOI: 10.1371/journal.pone.0014511
Source DB: PubMed Journal: PLoS One ISSN: 1932-6203 Impact factor: 3.240
The recovery rates of each of the CNV-calling programs depending on CNV length based on data of Kidd et al.
| # markers | # CNVs in the reference list from | # CNVs recovered by Birdsuite | # CNVs recovered by Partek | # CNVs recovered by PennCNV-Affy_trios | # CNVs recovered by PennCNV-Affy | # CNVs recovered by HelixTree |
| 1 | 329 | 6 (1.8%) | 0 | 0 | 0 | 2 (0.6%) |
| 2–5 | 249 | 71 (28.5%) | 0 | 3 (1.2%) | 2 (0.8%) | 19 (7.6%) |
| 6–10 | 112 | 47 (42.0%) | 10 (8.9%) | 20 (17.9%) | 11 (9.8%) | 28 (25%) |
| 10–20 | 73 | 32 (43.8%) | 26 (35.6%) | 27 (37.0%) | 24 (32.9%) | 17 (23.3%) |
| >20 | 130 | 91 (70.0%) | 70 (53.8%) | 76 (58.5%) | 72 (55.4%) | 54 (41.5%) |
*The pedigree information was incorporated with the calling of CNVs.
Recovery rate was calculated with the requirement of copy number consistency.
The recovery rates of each of the CNV-calling programs depending on CNV length based on data by Conrad et al. in 90 CEU samples.
| # markers | # CNVs in the reference list from Conrad | # CNVs recovered by Birdsuite | # CNVs recovered by Partek | # CNVs recovered by PennCNV-Affy_trios | # CNVs recovered by HelixTree |
| 1 | 28366 | 88 (0.31%) | 2 (0.007%) | 111 (0.39%) | 110 (0.39%) |
| 2–5 | 11837 | 1362 (11.51%) | 8 (0.068) | 138 (1.17%) | 486 (4.11%) |
| 6–10 | 3142 | 926 (29.47%) | 209 (6.65%) | 599 (19.06%) | 720 (22.92%) |
| 10–20 | 2754 | 973 (35.33%) | 507 (18.41%) | 747 (27.12%) | 711 (25.82%) |
| >20 | 5341 | 2547 (47.69%) | 1400 (26.21%) | 1883 (35.26%) | 1770 (33.14%) |
*5,341 CNVs spanned by >20 markers were included in this analysis.
**The pedigree information was incorporated with the calling of CNVs.
The average recovery rates of each of the CNV-calling programs depending on CNV frequency based on data by Conrad et al. in 90 CEU samples.
| Frequency(a) | # CNVs in the reference list from Conrad | # CNVs recovered by Birdsuite | # CNVs recovered by Partek | # CNVs recovered by PennCNV-Affy_trios | # CNVs recovered by HelixTree |
| a< = 20% | 669 | 537 (80.27%) | 439 (65.62%) | 575 (85.95%) | 528 (78.93%) |
| 20%<a< = 40% | 488 | 270 (55.33%) | 221 (45.29%) | 329 (67.41%) | 314 (64.34%) |
| 40%<a< = 60% | 793 | 579 (73.01%) | 210 (26.48) | 373 (47.04%) | 442 (55.74%) |
| 60%<a< = 80% | 765 | 352 (46.01%) | 105 (13.73%) | 139 (18.17%) | 123 (16.08%) |
| 80%<a< = 1 | 2626 | 801 (30.50%) | 284 (10.81%) | 366 (13.94%) | 342 (13.02%) |
*CNVs spanned by more than 20 markers were included in this analysis.
**Pedigree information was incorporated.
CNVs detected in 8 HapMap individuals shared between two studies.
| Size (kb) | # of CNVs in Conrad | # detected by Kidd | # of CNVs in Kidd | # detected by Conrad |
| ≤5 | 3024 | 38 (1.27%) | 441 | 1 (0.23%) |
| 5to 10 | 647 | 149 (23.02%) | 574 | 13 (2.26%) |
| 10 to 50 | 514 | 146 (28.40%) | 8174 | 300(3.67%) |
| 50 to 100 | 180 | 19 (10.56%) | 217 | 39 (17.97%) |
| 100 to 1000 | 172 | 11 (6.40%) | 107 | 11 (10.28%) |
*Criterion: Two CNVs from Conrad et al and Kidd et al were considered as the same if they shared at least 25% of the total length spanned. One CNV from Conrad et al can share with more than one CNV from Kidd et al, vice versa.
Figure 1The average number of CNVs per individual called by each of the four CNV-calling programs.
The average number of CNVs per individual varies greatly among programs, especially for CNVs less than 10 kb. The X axis represents length of CNVs. The Y axis is the average number of CNVs per individual.
The number of singleton CNVs called by each program, and percentage by which they overlap with calls made by the other programs.
| Program | Birdsuite | HelixTree | Partek | PennCNV-Affy | |
| Deletions | Shared | 306(90.5%) | 250(38.8%) | 279(90.0%) | 228(80.6%) |
| Program-specific | 32(9.5%) | 394(61.2%) | 31(10.0%) | 55(19.4%) | |
| Total | 338 | 644 | 310 | 283 | |
| Duplications | Shared | 401(74.7%) | 332(41.7%) | 354(90.8) | 289(85.3%) |
| Program-specific | 136(25.3%) | 465(58.3%) | 36(9.2%) | 50(14.7%) | |
| Total | 537 | 797 | 390 | 339 |
*Data format: number of events (percentage of events shared by other programs).
**Shared singleton deletions or duplications were defined as CNVs called by one program that overlapped at all with singleton deletions or duplications called by any other program.
***Program-specific CNVs are those that did not overlap at all with any singleton deletions or duplications called by any other program.
Positive predictive value for rare CNVs of each program, based on qPCR validation of their program-specific singletons.
| Programs | Birdsuite | HelixTree | Partek | PennCNV-Affy |
| Deletions | 5/5 = 100% | 0/5 = 0% | 5/5 = 100% | 3/5 = 60% |
| Duplications | 2/5 = 40% | 2/5 = 40% | 2/6 = 33.3% | 4/6 = 66.7% |
*Positive predictive value: true positive/(true positive + false positive). For each region, five samples were tested, one with a putative deletion/duplication, the other four with two putative).
Frequencies with which each program calls three common CNVs identified by Canary in the BiGS dataset.
| ID | Birdsuite | HelixTree | Partek | PennCNV-Affy | |
| CNP2157 | Frequency of Duplications | 1833(90.1%) | 29 (1.4%) | 24 (1.2%) | 15 (0.7%) |
| Frequency of Deletions | 1 (0.05%) | 187 (9.2%) | 42 (2.1%) | 38 (1.9%) | |
| CNP1293 | Frequency of Duplications | 1 (0.05%) | 664 (32.6%) | 619 (30.4%) | 508 (25.0%) |
| Frequency of Deletions | 1277(62.8%) | 449 (22.1%) | 344 (16.9%) | 262 (12.9%) | |
| CNP2057 | Frequency of Duplications | 170(8.4%) | 254 (12.5%) | 197 (9.7%) | 103 (5.1%) |
| Frequency of Deletions | 653(32.1%) | 248 (12.2%) | 142 (7.0%) | 188 (9.2%) |
qPCR validation of the calls made by each program for three common CNVs identified by Canary in the BiGS dataset.
| CNP | Birdsuite | HelixTree | Partek | PennCNV-Affy | |
| CNP2157 | False positive rate | 100% | 8.5% | 0.0% | 0.0% |
| False negative rate | 100% | 66.7% | 66.7% | 66.7% | |
| CNP1293 | False positive rate | 0.0% | 96.9% | 71.9% | 71.9% |
| False negative rate | 1.9% | 80.8% | 80.8% | 80.8% | |
| CNP2057 | False positive rate | 55.1% | 0.0% | 0.0% | 0.0% |
| False negative rate | 62.5% | 46.9% | 62.5% | 62.5% |
Settings used for each of the four software suites.
| Software | Plate-wise quantile normalization | Detecting algorithm | Parameters |
| Birdsuite | Yes (APT) | HMM | Using population-specific prior models |
| HelixTree | Yes (HelixTree) | Segmentation | Default |
| Partek | Yes (HelixTree) | Segmentation | Default |
| PennCNV-Affy | Yes (APT) | HMM | No prior models |
For HelixTree and Partek, default settings were used; normalization was done before CNV calling. For PennCNV-Affy, the standard procedure was followed without wave adjustment. For Birdsuite, population-specific prior models were employed.
Platewise normalization was done by Affy Power Tools (APT1.10.0) plug in Birdsuite/PennCNV-Affy.
Normalization was done by HelixTree.
For Canary, the appropriate prior model was selected based on the ancestry of the sample.