Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Biological sequence compression algorithms.

Literature DB >> 11700586

Biological sequence compression algorithms.

Abstract

Today, more and more DNA sequences are becoming available. The information about DNA sequences are stored in molecular biology databases. The size and importance of these databases will be bigger and bigger in the future, therefore this information must be stored or communicated efficiently. Furthermore, sequence compression can be used to define similarities between biological sequences. The standard compression algorithms such as gzip or compress cannot compress DNA sequences, but only expand them in size. On the other hand, CTW (Context Tree Weighting Method) can compress DNA sequences less than two bits per symbol. These algorithms do not use special structures of biological sequences. Two characteristic structures of DNA sequences are known. One is called palindromes or reverse complements and the other structure is approximate repeats. Several specific algorithms for DNA sequences that use these structures can compress them less than two bits per symbol. In this paper, we improve the CTW so that characteristic structures of DNA sequences are available. Before encoding the next symbol, the algorithm searches an approximate repeat and palindrome using hash and dynamic programming. If there is a palindrome or an approximate repeat with enough length then our algorithm represents it with length and distance. By using this preprocessing, a new program achieves a little higher compression ratio than that of existing DNA-oriented compression algorithms. We also describe new compression algorithm for protein sequences.

Mesh：

Year: 2000 PMID： 11700586

Source DB: PubMed Journal: Genome Inform Ser Workshop Genome Inform

Keyword Cloud
Cited

16 in total

Biological sequence compression algorithms.

1. Compressing proteomes: the relevance of medium range correlations.

2. Efficient storage of high throughput DNA sequencing data using reference-based compression.

3. Data Compression Concepts and Algorithms and their Applications to Bioinformatics.

4. An Adaptive Difference Distribution-based Coding with Hierarchical Tree Structure for DNA Sequence Compression.

5. Adaptive efficient compression of genomes.

6. Data structures and compression algorithms for high-throughput sequencing technologies.

7. Reference-based compression of short-read sequences using path encoding.

8. Cross chromosomal similarity for DNA sequence compression.

9. Compressing DNA sequence databases with coil.

10. DNA-COMPACT: DNA COMpression based on a pattern-aware contextual modeling technique.