Heng Li1. 1. Medical Population Genetics Program, Broad Institute of Harvard and MIT, Cambridge, MA 02142, USA.
Abstract
MOTIVATION: Whole-genome high-coverage sequencing has been widely used for personal and cancer genomics as well as in various research areas. However, in the lack of an unbiased whole-genome truth set, the global error rate of variant calls and the leading causal artifacts still remain unclear even given the great efforts in the evaluation of variant calling methods. RESULTS: We made 10 single nucleotide polymorphism and INDEL call sets with two read mappers and five variant callers, both on a haploid human genome and a diploid genome at a similar coverage. By investigating false heterozygous calls in the haploid genome, we identified the erroneous realignment in low-complexity regions and the incomplete reference genome with respect to the sample as the two major sources of errors, which press for continued improvements in these two areas. We estimated that the error rate of raw genotype calls is as high as 1 in 10-15 kb, but the error rate of post-filtered calls is reduced to 1 in 100-200 kb without significant compromise on the sensitivity. AVAILABILITY AND IMPLEMENTATION: BWA-MEM alignment and raw variant calls are available at http://bit.ly/1g8XqRt scripts and miscellaneous data at https://github.com/lh3/varcmp. CONTACT: hengli@broadinstitute.org SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
MOTIVATION: Whole-genome high-coverage sequencing has been widely used for personal and cancer genomics as well as in various research areas. However, in the lack of an unbiased whole-genome truth set, the global error rate of variant calls and the leading causal artifacts still remain unclear even given the great efforts in the evaluation of variant calling methods. RESULTS: We made 10 single nucleotide polymorphism and INDEL call sets with two read mappers and five variant callers, both on a haploid human genome and a diploid genome at a similar coverage. By investigating false heterozygous calls in the haploid genome, we identified the erroneous realignment in low-complexity regions and the incomplete reference genome with respect to the sample as the two major sources of errors, which press for continued improvements in these two areas. We estimated that the error rate of raw genotype calls is as high as 1 in 10-15 kb, but the error rate of post-filtered calls is reduced to 1 in 100-200 kb without significant compromise on the sensitivity. AVAILABILITY AND IMPLEMENTATION: BWA-MEM alignment and raw variant calls are available at http://bit.ly/1g8XqRt scripts and miscellaneous data at https://github.com/lh3/varcmp. CONTACT: hengli@broadinstitute.org SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Authors: Samuel Levy; Granger Sutton; Pauline C Ng; Lars Feuk; Aaron L Halpern; Brian P Walenz; Nelson Axelrod; Jiaqi Huang; Ewen F Kirkness; Gennady Denisov; Yuan Lin; Jeffrey R MacDonald; Andy Wing Chun Pang; Mary Shago; Timothy B Stockwell; Alexia Tsiamouri; Vineet Bafna; Vikas Bansal; Saul A Kravitz; Dana A Busam; Karen Y Beeson; Tina C McIntosh; Karin A Remington; Josep F Abril; John Gill; Jon Borman; Yu-Hui Rogers; Marvin E Frazier; Stephen W Scherer; Robert L Strausberg; J Craig Venter Journal: PLoS Biol Date: 2007-09-04 Impact factor: 8.029
Authors: David L Goode; Sally M Hunter; Maria A Doyle; Tao Ma; Simone M Rowley; David Choong; Georgina L Ryland; Ian G Campbell Journal: Genome Med Date: 2013-09-30 Impact factor: 11.117
Authors: Goncalo R Abecasis; Adam Auton; Lisa D Brooks; Mark A DePristo; Richard M Durbin; Robert E Handsaker; Hyun Min Kang; Gabor T Marth; Gil A McVean Journal: Nature Date: 2012-11-01 Impact factor: 49.962
Authors: Nicola D Roberts; R Daniel Kortschak; Wendy T Parker; Andreas W Schreiber; Susan Branford; Hamish S Scott; Garique Glonek; David L Adelson Journal: Bioinformatics Date: 2013-07-09 Impact factor: 6.937
Authors: Olivier Harismendy; Pauline C Ng; Robert L Strausberg; Xiaoyun Wang; Timothy B Stockwell; Karen Y Beeson; Nicholas J Schork; Sarah S Murray; Eric J Topol; Samuel Levy; Kelly A Frazer Journal: Genome Biol Date: 2009-03-27 Impact factor: 13.583
Authors: Klaus Schmitz-Abe; Qifei Li; Samantha M Rosen; Neeharika Nori; Jill A Madden; Casie A Genetti; Monica H Wojcik; Sadhana Ponnaluri; Cynthia S Gubbels; Jonathan D Picker; Anne H O'Donnell-Luria; Timothy W Yu; Olaf Bodamer; Catherine A Brownstein; Alan H Beggs; Pankaj B Agrawal Journal: Eur J Hum Genet Date: 2019-04-12 Impact factor: 4.246
Authors: Gregg W C Thomas; Richard J Wang; Arthi Puri; R Alan Harris; Muthuswamy Raveendran; Daniel S T Hughes; Shwetha C Murali; Lawrence E Williams; Harsha Doddapaneni; Donna M Muzny; Richard A Gibbs; Christian R Abee; Mary R Galinski; Kim C Worley; Jeffrey Rogers; Predrag Radivojac; Matthew W Hahn Journal: Curr Biol Date: 2018-09-27 Impact factor: 10.834