Literature DB >> 25064564

Compression and fast retrieval of SNP data.

Francesco Sambo1, Barbara Di Camillo1, Gianna Toffolo1, Claudio Cobelli1.   

Abstract

MOTIVATION: The increasing interest in rare genetic variants and epistatic genetic effects on complex phenotypic traits is currently pushing genome-wide association study design towards datasets of increasing size, both in the number of studied subjects and in the number of genotyped single nucleotide polymorphisms (SNPs). This, in turn, is leading to a compelling need for new methods for compression and fast retrieval of SNP data.
RESULTS: We present a novel algorithm and file format for compressing and retrieving SNP data, specifically designed for large-scale association studies. Our algorithm is based on two main ideas: (i) compress linkage disequilibrium blocks in terms of differences with a reference SNP and (ii) compress reference SNPs exploiting information on their call rate and minor allele frequency. Tested on two SNP datasets and compared with several state-of-the-art software tools, our compression algorithm is shown to be competitive in terms of compression rate and to outperform all tools in terms of time to load compressed data.
AVAILABILITY AND IMPLEMENTATION: Our compression and decompression algorithms are implemented in a C++ library, are released under the GNU General Public License and are freely downloadable from http://www.dei.unipd.it/~sambofra/snpack.html.
© The Author 2014. Published by Oxford University Press. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com.

Mesh:

Year:  2014        PMID: 25064564      PMCID: PMC4609015          DOI: 10.1093/bioinformatics/btu495

Source DB:  PubMed          Journal:  Bioinformatics        ISSN: 1367-4803            Impact factor:   6.937


  16 in total

1.  PLINK: a tool set for whole-genome association and population-based linkage analyses.

Authors:  Shaun Purcell; Benjamin Neale; Kathe Todd-Brown; Lori Thomas; Manuel A R Ferreira; David Bender; Julian Maller; Pamela Sklar; Paul I W de Bakker; Mark J Daly; Pak C Sham
Journal:  Am J Hum Genet       Date:  2007-07-25       Impact factor: 11.025

Review 2.  Bringing genome-wide association findings into clinical use.

Authors:  Teri A Manolio
Journal:  Nat Rev Genet       Date:  2013-07-09       Impact factor: 53.242

3.  Genome compression: a novel approach for large collections.

Authors:  Sebastian Deorowicz; Agnieszka Danek; Szymon Grabowski
Journal:  Bioinformatics       Date:  2013-08-21       Impact factor: 6.937

Review 4.  Rare and common variants: twenty arguments.

Authors:  Greg Gibson
Journal:  Nat Rev Genet       Date:  2012-01-18       Impact factor: 53.242

Review 5.  Haplotype blocks and linkage disequilibrium in the human genome.

Authors:  Jeffrey D Wall; Jonathan K Pritchard
Journal:  Nat Rev Genet       Date:  2003-08       Impact factor: 53.242

6.  Fast and accurate genotype imputation in genome-wide association studies through pre-phasing.

Authors:  Bryan Howie; Christian Fuchsberger; Matthew Stephens; Jonathan Marchini; Gonçalo R Abecasis
Journal:  Nat Genet       Date:  2012-07-22       Impact factor: 38.330

7.  Handling the data management needs of high-throughput sequencing data: SpeedGene, a compression algorithm for the efficient storage of genetic data.

Authors:  Dandi Qiao; Wai-Ki Yip; Christoph Lange
Journal:  BMC Bioinformatics       Date:  2012-05-16       Impact factor: 3.169

8.  Efficient haplotype matching and storage using the positional Burrows-Wheeler transform (PBWT).

Authors:  Richard Durbin
Journal:  Bioinformatics       Date:  2014-01-09       Impact factor: 6.937

9.  Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls.

Authors: 
Journal:  Nature       Date:  2007-06-07       Impact factor: 49.962

10.  An integrated map of genetic variation from 1,092 human genomes.

Authors:  Goncalo R Abecasis; Adam Auton; Lisa D Brooks; Mark A DePristo; Richard M Durbin; Robert E Handsaker; Hyun Min Kang; Gabor T Marth; Gil A McVean
Journal:  Nature       Date:  2012-11-01       Impact factor: 49.962

View more
  2 in total

1.  Second-generation PLINK: rising to the challenge of larger and richer datasets.

Authors:  Christopher C Chang; Carson C Chow; Laurent Cam Tellier; Shashaank Vattikuti; Shaun M Purcell; James J Lee
Journal:  Gigascience       Date:  2015-02-25       Impact factor: 6.524

2.  Efficiently Summarizing Relationships in Large Samples: A General Duality Between Statistics of Genealogies and Genomes.

Authors:  Peter Ralph; Kevin Thornton; Jerome Kelleher
Journal:  Genetics       Date:  2020-05-01       Impact factor: 4.562

  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.