| Literature DB >> 22369457 |
Yoichi Takenaka1, Shigeto Seno, Hideo Matsuda.
Abstract
BACKGROUND: With the advent of next-generation sequencers, the growing demands to map short DNA sequences to a genome have promoted the development of fast algorithms and tools. The tools commonly used today are based on either a hash table or the suffix array/Burrow-Wheeler transform. These algorithms are the best suited to finding the genome position of exactly matching short reads. However, they have limited capacity to handle the mismatches. To find n-mismatches, they requiresEntities:
Mesh:
Year: 2011 PMID: 22369457 PMCID: PMC3333191 DOI: 10.1186/1471-2164-12-S3-S8
Source DB: PubMed Journal: BMC Genomics ISSN: 1471-2164 Impact factor: 3.969
Figure 1Hash tables for three methods. Three methods to find genome positions of 1-mismatch from the subsequence AAGT. Genome position 1000 is ACGT, which is the 1-mismatch of the subsequence. The first method refers to the hash table 16 times. The second method refers to the table just once, but the table is 16-fold larger. The third method refers to the table three times. After getting position 1002 from the hash table, the method elongates the alignment toward the front of the sequence.
Figure 2Relationship among 16 subsequences. Graphical depiction of subsequence “AAAAA” and 15 adjacent subsequences. Each node describes a subsequence and each edge indicates that the terminal nodes are of one nucleotide difference. The 15 subsequences are divided into five groups according to the position of the different nucleotide.
Figure 31-mismatch sequences of a non-code word sequence, “CAAAA”. Idea behind finding 1-mismatch sequences of the sequence “CAAAA”. The two circles indicate the equivalence classes. The sequence “CAAAA” belongs to an equivalence class whose center is “AAAAA” that holds 3 of the 15 1-mismatch sequences. Two of the other sequences, “CAAAT” and “CAAGA”, belong to a equivalence class whose center is “CAAGT”.
Addition and multiplication on GF(22)
| + | 0 | 1 | × | 0 | 1 | |||||
| 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | |||
| 1 | 1 | 0 | 1 | 0 | 1 | |||||
| 0 | 1 | 0 | 1 | |||||||
| 1 | 0 | 0 | 1 | |||||||
Figure 4Proposed hash table. Entries of proposed hash tables. Subsequence “AAAAA” is used as the key for storing the genome positions of “ACAAA” and “AAATA” because“AAAAA” is the center of the equivalence class that “ACAA” and“AAAT” belong to. The left hash table uses the center subsequence itself as a hash key. The right one uses the short code of the center subsequence as a hash key, where the short code is the information digits of the code. Short codes are described in Section Perfect Hamming Code.
Summary of our methods for lengths 5, 21, and 10 to refer to 1- and 2-mismatch and 1- and 2-gap sequences
| length | condition | #keys | #words | ratio | ||
|---|---|---|---|---|---|---|
| 5 | 1-mismatch | 6.625 | 16 | 41.4% | 1 + 15 | 1 + 15 |
| 2-mismatches | 27.25 | 106 | 25.7% | 1 + 15 + 90 | 1 + 15 + 90 | |
| 1-gap | 3.25 | 4 | 81.3% | 4 + 12 | 4 + 60 | |
| 2-gaps | 10 | 16 | 62.5% | 16 + 36 | – ∗1 | |
| 21 | 1-mismatch | 30.53 | 64 | 47.7% | 1 + 63 | 1 + 63 |
| 2-mismatches | 611.31 | 1954 | 31.3% | 1 + 63 | 1 + 63 | |
| 1-gap | 3.81 | 4 | 95.3% | 4 + 60 | 4 + 252 | |
| 2-gaps | 13.87 | 16 | 86.7% | 16 + 84 | 16 + 48 | |
| 10: Serialize | 1-mismatch | 12.25 | 31 | 39.5% | 1 + 30 | 1 + 30 |
| 10: Parallelize | 1-mismatch | 13.25 | 31 | 44.1% | 1 + 30 | 1 + 30 |
∗1 :s always includes one code word. ∗2: neither the first half nor the second half are code words. The reference formula when one of the two halves is a code word is 1 + 30x2 + 267x2 + 684x3 + 810x4. ∗3: neither the first half or second half are code words. The reference formula when one of the two halves is a code word is 1 + 30x2 + 42x2 + 54x3.
Figure 5Fifteen 1-mismatch subsequences of . Nodes represent subsequences and edges indicate a Hamming distance between two nodes of 1, namely, the relation of 1-mismatch. Edge labels indicate the position of the different digit (nucleotide). s belongs to a equivalence class of c(s), which also contains three other words. The other 12 words belong to six equivalence classes.
Figure 6Two-mismatch subsequences of . Dark-gray nodes are 2-mismatches whose Hamming distance from s is 2. Each belongs to an equivalence class that has other two 2-mismatche subsequences.
Figure 7Two-mismatch subsequences of . The equivalence classes they belong to are classified into four types.
Figure 8Three ways to elongate the length of subsequences.