Literature DB >> 10222406

Transformation distances: a family of dissimilarity measures based on movements of segments.

J S Varré1, J P Delahaye, E Rivals.   

Abstract

MOTIVATION: Evolution acts in several ways on DNA: either by mutating a base, or by inserting, deleting or copying a segment of the sequence (Ruddle, 1997; Russell, 1994; Li and Grauer, 1991). Classical alignment methods deal with point mutations (Waterman, 1995), genome-level mutations are studied using genome rearrangement distances (Bafna and Pevzner, 1993, 1995; Kececioglu and Sankoff, 1994; Kececioglu and Ravi, 1995). The latter distances generally operate, not on the sequences, but on an ordered list of genes. To our knowledge, no measure of distance attempts to compare sequences using a general set of segment-based operations.
RESULTS: Here we define a new family of distances, called transformation distances, which quantify the dissimilarity between two sequences in terms of segment-based events. We focus on the case where segment-copy, -reverse-copy and -insertion are allowed in our set of operations. Those events are weighted by their description length, but other sets of weights are possible when biological information is available. The transformation distance from sequence S to sequence T is then the Minimum Description Length among all possible scripts that build T knowing S with segment-based operations. The underlying idea is related to Kolmogorov complexity theory. We present an algorithm which, given two sequences S and T, computes exactly and efficiently the transformation distance from S to T. Unlike alignment methods, the method we propose does not necessarily respect the order of the residues within the compared sequences and is therefore able to account for duplications and translocations that cannot be properly described by sequence alignment. A biological application on Tnt1 tobacco retrotransposon is presented. AVAILABILITY: The algorithm and the graphical interface can be downloaded at http://www.lifl.fr/ approximately varre/TD

Entities:  

Mesh:

Substances:

Year:  1999        PMID: 10222406     DOI: 10.1093/bioinformatics/15.3.194

Source DB:  PubMed          Journal:  Bioinformatics        ISSN: 1367-4803            Impact factor:   6.937


  7 in total

1.  A genomic distance based on MUM indicates discontinuity between most bacterial species and genera.

Authors:  Marc Deloger; Meriem El Karoui; Marie-Agnès Petit
Journal:  J Bacteriol       Date:  2008-10-31       Impact factor: 3.490

2.  Shortest Path Edit Distance for Enhancing UMLS Integration and Audit.

Authors:  Alex Rudniy; James Geller; Min Song
Journal:  AMIA Annu Symp Proc       Date:  2010-11-13

3.  Recursive organizer (ROR): an analytic framework for sequence-based association analysis.

Authors:  Lue Ping Zhao; Xin Huang
Journal:  Hum Genet       Date:  2013-03-14       Impact factor: 4.132

4.  Mapping sequences by parts.

Authors:  Gilles Didier; Carito Guziolowski
Journal:  Algorithms Mol Biol       Date:  2007-09-19       Impact factor: 1.405

5.  Sequence alignment, mutual information, and dissimilarity measures for constructing phylogenies.

Authors:  Orion Penner; Peter Grassberger; Maya Paczuski
Journal:  PLoS One       Date:  2011-01-04       Impact factor: 3.240

6.  Evolutionary hierarchies of conserved blocks in 5'-noncoding sequences of dicot rbcS genes.

Authors:  Katie E Weeks; Nadia A Chuzhanova; Iain S Donnison; Ian M Scott
Journal:  BMC Evol Biol       Date:  2007-04-02       Impact factor: 3.260

7.  CLUSS: clustering of protein sequences based on a new similarity measure.

Authors:  Abdellali Kelil; Shengrui Wang; Ryszard Brzezinski; Alain Fleury
Journal:  BMC Bioinformatics       Date:  2007-08-04       Impact factor: 3.169

  7 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.