Literature DB >> 16108721

Designing multiple simultaneous seeds for DNA similarity search.

Yanni Sun1, Jeremy Buhler.   

Abstract

The challenge of similarity search in massive DNA sequence databases has inspired major changes in BLAST-style alignment tools, which accelerate search by inspecting only pairs of sequences sharing a common short "seed," or pattern of matching residues. Some of these changes raise the possibility of improving search performance by probing sequence pairs with several distinct seeds, any one of which is sufficient for a seed match. However, designing a set of seeds to maximize their combined sensitivity to biologically meaningful sequence alignments is computationally difficult, even given recent advances in designing single seeds. This work describes algorithmic improvements to seed design that address the problem of designing a set of n seeds to be used simultaneously. We give a new local search method to optimize the sensitivity of seed sets. The method relies on efficient incremental computation of the probability that an alignment contains a match to a seed pi, given that it has already failed to match any of the seeds in a set Pi. We demonstrate experimentally that multi-seed designs, even with relatively few seeds, can be significantly more sensitive than even optimized single-seed designs.

Mesh:

Year:  2005        PMID: 16108721     DOI: 10.1089/cmb.2005.12.847

Source DB:  PubMed          Journal:  J Comput Biol        ISSN: 1066-5277            Impact factor:   1.479


  10 in total

1.  Graemlin: general and robust alignment of multiple large interaction networks.

Authors:  Jason Flannick; Antal Novak; Balaji S Srinivasan; Harley H McAdams; Serafim Batzoglou
Journal:  Genome Res       Date:  2006-08-09       Impact factor: 9.043

2.  Short Read Mapping: An Algorithmic Tour.

Authors:  Stefan Canzar; Steven L Salzberg
Journal:  Proc IEEE Inst Electr Electron Eng       Date:  2015-09-07       Impact factor: 10.961

3.  Global, highly specific and fast filtering of alignment seeds.

Authors:  Matthis Ebel; Giovanna Migliorelli; Mario Stanke
Journal:  BMC Bioinformatics       Date:  2022-06-10       Impact factor: 3.307

4.  Designing Efficient Spaced Seeds for SOLiD Read Mapping.

Authors:  Laurent Noé; Marta Gîrdea; Gregory Kucherov
Journal:  Adv Bioinformatics       Date:  2010-09-16

5.  Hit integration for identifying optimal spaced seeds.

Authors:  Won-Hyoung Chung; Seong-Bae Park
Journal:  BMC Bioinformatics       Date:  2010-01-18       Impact factor: 3.169

6.  Best hits of 11110110111: model-free selection and parameter-free sensitivity calculation of spaced seeds.

Authors:  Laurent Noé
Journal:  Algorithms Mol Biol       Date:  2017-02-14       Impact factor: 1.405

7.  Effective sequence similarity detection with strobemers.

Authors:  Kristoffer Sahlin
Journal:  Genome Res       Date:  2021-10-19       Impact factor: 9.043

8.  BFAST: an alignment tool for large scale genome resequencing.

Authors:  Nils Homer; Barry Merriman; Stanley F Nelson
Journal:  PLoS One       Date:  2009-11-11       Impact factor: 3.240

9.  PerM: efficient mapping of short sequencing reads with periodic full sensitive spaced seeds.

Authors:  Yangho Chen; Tade Souaiaia; Ting Chen
Journal:  Bioinformatics       Date:  2009-08-12       Impact factor: 6.937

10.  Universal seeds for cDNA-to-genome comparison.

Authors:  Leming Zhou; Jonathan Stanton; Liliana Florea
Journal:  BMC Bioinformatics       Date:  2008-01-23       Impact factor: 3.169

  10 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.