| Literature DB >> 22974120 |
Xiao Yang1, Patrick Charlebois, Sante Gnerre, Matthew G Coole, Niall J Lennon, Joshua Z Levin, James Qu, Elizabeth M Ryan, Michael C Zody, Matthew R Henn.
Abstract
BACKGROUND: Extensive genetic diversity in viral populations within infected hosts and the divergence of variants from existing reference genomes impede the analysis of deep viral sequencing data. A de novo population consensus assembly is valuable both as a single linear representation of the population and as a backbone on which intra-host variants can be accurately mapped. The availability of consensus assemblies and robustly mapped variants are crucial to the genetic study of viral disease progression, transmission dynamics, and viral evolution. Existing de novo assembly techniques fail to robustly assemble ultra-deep sequence data from genetically heterogeneous populations such as viruses into full-length genomes due to the presence of extensive genetic variability, contaminants, and variable sequence coverage.Entities:
Mesh:
Year: 2012 PMID: 22974120 PMCID: PMC3469330 DOI: 10.1186/1471-2164-13-475
Source DB: PubMed Journal: BMC Genomics ISSN: 1471-2164 Impact factor: 3.969
Figure 1Schematic of the assembly algorithm.
Assembly Results for , and 454
| | | | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| | V4526 | 10 | 100 | 1 | 100 | 95.31 | 0 | 1/11 | 248 | 0.42 | |
| | | 19 | 34.51 | 18 | 4.29 | 16.31 | 17.23 | -a | 79 | 6.40 | |
| | | 4 | 100 | 1 | 100 | 94.69 | 0 | 5/11 | 507 | 1.10 | |
| | V4528 | 9 | 100 | 1 | 100 | 95.01 | 0 | 1/11 | 305 | 0.44 | |
| | | 24 | 39.26 | 22 | 3.75 | -a | -a | - | 79 | 6.20 | |
| WNV | | 3 | 100 | 1 | 100 | 94.52 | 0 | 2/11 | 379 | 0.12 | |
| | V5044 | 8 | 100 | 1 | 100 | 95.08 | 0 | 0/11 | 441 | 0.59 | |
| | | 32 | 40.23 | 28 | 3.16 | 8.84 | 17.45 | - | 117 | 6.40 | |
| | | 5 | 100 | 1 | 100 | 94.22 | 0.01 | 5/11 | 387 | 0.21 | |
| | V5048 | 9 | 100 | 1 | 100 | 95.08 | 0 | 0/11 | 212 | 0.43 | |
| | | 17 | 24.52 | 15 | 3.49 | 5.67 | 15.90 | - | 90 | 6.40 | |
| | | 1 | 100 | 1 | 100 | 95.08 | 0 | 1/11 | 453 | 0.14 | |
| | V4809 | 6 | 100 | 1 | 100 | 95.64 | 0.01 | 0/11 | 510 | 0.92 | |
| | | 40 | 59.7 | 33 | 3.63 | 16.05 | 17.78 | - | 184 | 6.50 | |
| | | 7 | 100 | 1 | 100 | 94.6 | 0.019 | 5/11 | 474 | 0.27 | |
| | V4813 | 12 | 100 | 1 | 100 | 95.12 | 0.01 | 0/11 | 669 | 1.02 | |
| | | 49 | 64.33 | 40 | 3.66 | 18.21 | 18.54 | - | 193 | 6.50 | |
| DENV | | 2 | 100 | 1 | 100 | 94.18 | 0.04 | 7/11 | 492 | 0.55 | |
| | V4816 | 9 | 100 | 1 | 100 | 95.52 | 0 | 0/11 | 677 | 0.91 | |
| | | 37 | 53.85 | 31 | 3.76 | 11.91 | 16.63 | - | 167 | 6.50 | |
| | | 5 | 100 | 2 | 100 | 94.84 | 0.20 | 5/11 | 471 | 0.32 | |
| | V4820 | 14 | 100 | 2 | 82.45 | 93.46 | 0 | 0/11 | 1158 | 1.20 | |
| | | 56 | 70.73 | 46 | 3.62 | 13.37 | 17.20 | - | 234 | 6.50 | |
| | | 13 | 100 | 2 | 76.68 | 91.59 | 0.18 | 6/11 | 462 | 0.17 | |
| | V5937 | 12 | 100 | 2 | 93.58 | 93.8 | 0.02 | 0/9 | 516 | 0.86 | |
| | | 42 | 48.01 | 30 | 3.95 | 17.72 | 17.71 | - | 142 | 6.40 | |
| | | 16 | 100 | 1 | 100 | 86.15 | 0.55 | 6/9 | 406 | 0.15 | |
| | V5938 | 18 | 100 | 1 | 100 | 93.69 | 0.01 | 0/9 | 281 | 0.62 | |
| | | 28 | 40.5 | 25 | 4.16 | 11.33 | 16.08 | - | 111 | 6.40 | |
| HIV | | 15 | 100 | 1 | 100 | 88.01 | 0.43 | 5/9 | 443 | 0.21 | |
| | V5943 | 9 | 100 | 1 | 100 | 95.58 | 0.05 | 1/9 | 96.5 | 0.21 | |
| | | 24 | 32.03 | 19 | 4.1 | 12.15 | 18.37 | - | 40 | 6.50 | |
| | | 9 | 97.16 | 1 | 97.16 | 92.52 | 0.80 | 4/9 | 583 | 0.55 | |
| | V5945 | 13 | 100 | 2 | 98.83 | 94.44 | 0.09 | 0/9 | 576 | 0.60 | |
| | | 31 | 49.02 | 25 | 4.21 | 13.53 | 16.85 | - | 110 | 6.20 | |
| 7 | 100 | 2 | 83.32 | 89.54 | 1.18 | 4/9 | 465 | 0.17 | |||
aSOAPdenovo assembly is highly fragmented, a large number of short contigs were merged using the reference genome, leading to the inclusion of many low frequency variants that considerably increased the percentage of non-dominant bases found in the assembly. In the case of sample V4528, Mosaik failed to report read alignment to the consensus. The number of genes with frame shift is not measured for SOAPdenovo. For run time and memory, †Soapdenovo uses 8 threads while the other two use 1 thread. AV454 is run on a subset of the reads (∼ 11k, equivalent to 1% – 8% of input).
Figure 2Coverage plot. Fold sequence coverage across the target regions of one representative sample for DENV, WNV, and HIV full-length genomes. Coverage is measured as the total number of reads uniquely aligning over a given residue; alignments are to standard references (see Methods).