| Literature DB >> 23300661 |
Robert Lindner1, Caroline C Friedel.
Abstract
Transcriptome sequencing (RNA-Seq) overcomes limitations of previously used RNA quantification methods and provides one experimental framework for both high-throughput characterization and quantification of transcripts at the nucleotide level. The first step and a major challenge in the analysis of such experiments is the mapping of sequencing reads to a transcriptomic origin including the identification of splicing events. In recent years, a large number of such mapping algorithms have been developed, all of which have in common that they require algorithms for aligning a vast number of reads to genomic or transcriptomic sequences. Although the FM-index based aligner Bowtie has become a de facto standard within mapping pipelines, a much larger number of possible alignment algorithms have been developed also including other variants of FM-index based aligners. Accordingly, developers and users of RNA-seq mapping pipelines have the choice among a large number of available alignment algorithms. To provide guidance in the choice of alignment algorithms for these purposes, we evaluated the performance of 14 widely used alignment programs from three different algorithmic classes: algorithms using either hashing of the reference transcriptome, hashing of reads, or a compressed FM-index representation of the genome. Here, special emphasis was placed on both precision and recall and the performance for different read lengths and numbers of mismatches and indels in a read. Our results clearly showed the significant reduction in memory footprint and runtime provided by FM-index based aligners at a precision and recall comparable to the best hash table based aligners. Furthermore, the recently developed Bowtie 2 alignment algorithm shows a remarkable tolerance to both sequencing errors and indels, thus, essentially making hash-based aligners obsolete.Entities:
Mesh:
Year: 2012 PMID: 23300661 PMCID: PMC3530550 DOI: 10.1371/journal.pone.0052403
Source DB: PubMed Journal: PLoS One ISSN: 1932-6203 Impact factor: 3.240
Alignment software undergoing complete evaluation.
| Algorithmic Class | Aligner | Version |
|
| ||
|
| MAQ | 0.7.1 |
| RazerS | 1.1 | |
| RMAP | 2.05 | |
|
| BFAST | 0.6.4e |
| Genomemapper | 0.4.3s | |
| GNUMAP | 2.2.3 | |
| Mosaik | 0.7.1 | |
| mrFast | 2.0.0.5 | |
| Novoalign | 2.07.06 | |
| SHRiMP | 2.1.1 | |
|
| ||
|
| Bowtie | 0.12.7 |
| Bowtie 2 | 2.0.0-beta7 | |
| BWA | 0.5.9-r16 | |
| SOAP2 | 2.21 | |
Popular alignment tools for whole-genome applications were chosen to represent the major algorithmic classes.
Figure 1Workflow of alignment evaluation.
Reads were simulated from chromosome 21 of the ENSEMBL human transcriptome using dwgsim and aligned to the transcriptome with various parameter combinations for each alignment tool. The best parameter combination for each aligner was selected based on -measure first and runtime and memory consumption second in case of ties. The best parameters were then used to evaluate the alignment algorithms on larger test sets simulated from chromosomes 1-22, excluding chromosome 21.
Parameter Sensitivity.
| Aligner | Parameters Tested | DF |
|
| Read Indexing | |
|
| 145 | 0.05 |
|
| 162 | 0.003 |
|
| 48 | 0.1 |
| Reference Indexing | ||
|
| 6 | 0.40 |
|
| 216 | 0.07 |
|
| 324 | 0.03 |
|
| 108 | 0.01 |
|
| 7 | 1.07 |
|
| 288 | 0.22 |
|
| 432 | 0.06 |
|
| FM-Index Based | |
|
| 180 | 0.04 |
|
| 864 | 0.004 |
|
| 16 | 0.04 |
|
| 72 | 0.37 |
Dispersion of the -measure across all parameter settings tested for each alignment algorithm is shown for the 72 bp read length training set from human chromosome 21. describes the sensitivity of an alignment algorithm to parameter changes. For BFAST and Mosaik, not all available parameters were investigated due to the modular structure of the application.
Figure 2Influence of parameter selection.
Precision (x-axis) and recall (y-axis) are shown for SOAP2 (black triangles) and Bowtie 2 (gray diamonds) alignment tools at a read length of 72 bp. Different parameters had only little impact on the alignment performance of Bowtie 2 whereas the recall of SOAP2 was scattered widely. This is reflected by the high -measure dispersion value of SOAP2.
Evaluation on test set.
| Aligner | Reads | Precision | Recall | F | Memory | Runtime |
|
| Read Indexing | |||||
| 36 | 0.913 | 0.908 | 0.911 | 3.1G | 2∶20 h | |
|
| 72 | 0.922 | 0.918 | 0.920 | 3.4G | 2∶20 h |
| 100 | 0.925 | 0.920 | 0.923 | 3.7G | 2∶30 h | |
| 36 | 0.883 | 0.980 | 0.929 | 8.8G | 77∶05 h | |
|
| 72 | – | – | – | – | – |
| 100 | – | – | – | – | – | |
| 36 | 0.912 | 0.483 | 0.631 | 5.4G | 18∶00 h | |
|
| 72 | 0.919 | 0.503 | 0.650 | 6.0G | 10∶30 h |
| 100 | 0.920 | 0.524 | 0.668 | 6.0G | 7∶05 h | |
| Reference Indexing | ||||||
| 36 | 0.898 | 0.577 | 0.703 | 5.2G | 2∶20 h | |
|
| 72 | 0.898 | 0.801 | 0.847 | 6.0G | 9∶30 h |
| 100 | 0.910 | 0.819 | 0.862 | 6.3G | 18∶10 h | |
| 36 | 0.853 | 0.953 | 0.900 | 3.7G | 3∶10 h | |
|
| 72 | 0.870 | 0.967 | 0.916 | 3.8G | 32∶40 h |
| 100 | 0.877 | 0.971 | 0.921 | 3.8G | 20∶10 h | |
| 36 | – | – | – | – | – | |
|
| 72 | – | – | – | – | – |
| 100 | – | – | – | – | – | |
| 36 | – | – | – | – | – | |
|
| 72 | – | – | – | – | – |
| 100 | – | – | – | – | – | |
| 36 | 0.914 | 0.914 | 0.914 | 2.6G | 1∶45 h | |
|
| 72 | 0.921 | 0.919 | 0.920 | 2.6G | 3∶20 h |
| 100 | 0.924 | 0.923 | 0.924 | 2.6G | 4∶20 h | |
| 36 | – | – | – | – | – | |
|
| 72 | – | – | – | – | – |
| 100 | – | – | – | – | – | |
|
| FM-Index Based | |||||
| 36 | 0.908 | 0.895 | 0.902 | 750M | 0∶30 h | |
|
| 72 | 0.919 | 0.896 | 0.907 | 750M | 0∶30 h |
| 100 | 0.923 | 0.874 | 0.898 | 750M | 0∶35 h | |
| 36 | 0.912 | 0.911 | 0.912 | 1.3G | 2∶10 h | |
|
| 72 | 0.920 | 0.919 | 0.920 | 1.3G | 3∶20 h |
| 100 | 0.923 | 0.922 | 0.923 | 1.3G | 2∶30 h | |
| 36 | 0.912 | 0.875 | 0.893 | 1.3G | 0∶35 h | |
|
| 72 | 0.921 | 0.915 | 0.918 | 1.5G | 1∶10 h |
| 100 | 0.955 | 0.955 | 0.955 | 1.6G | 1∶00 h | |
| 36 | 0.891 | 0.853 | 0.872 | 2.9G | 0∶15 h | |
|
| 72 | 0.886 | 0.751 | 0.812 | 2.9G | 1∶00 h |
| 100 | 0.907 | 0.834 | 0.869 | 2.9G | 1∶00 h |
Evaluation of optimal alignment parameters on 5 million reads from chromosomes 1-22, excluding 21. Runtime and memory caps were set at 10 CPU days and 16G, respectively. Processes exceeding these limits were killed and partial results were not evaluated. Superscripts indicate the following:
Runtime and memory consumption could not be recorded accurately because of the modular structure of the application which makes the automated evaluation of all parameter combinations difficult;
Supports only single-end alignment;
Exceeded the memory cap of 16G;
Exceeded the runtime cap of 10 CPU days.
Figure 3Sensitivity of alignment algorithms to errors and indels.
The dependence of alignment precision (dotted lines ), recall (dashed lines ) and -measure (solid lines –) on the number of sequencing errors (gray) and indels (black) was evaluated for each alignment algorithm analyzed in this study. In this figure, results are shown for the 72 bp long reads from the test set (5 million reads), excluding those algorithms which exceeded the memory and runtime cap. Alignment algorithms were sorted by algorithmic design: (A) Read indexing based aligners, (B) Reference indexing based aligners, (C) FM-Index based aligners.