| Literature DB >> 17725821 |
Rodrigo Gouveia-Oliveira1, Peter W Sackett, Anders G Pedersen.
Abstract
BACKGROUND: The presence of gaps in an alignment of nucleotide or protein sequences is often an inconvenience for bioinformatical studies. In phylogenetic and other analyses, for instance, gapped columns are often discarded entirely from the alignment.Entities:
Mesh:
Year: 2007 PMID: 17725821 PMCID: PMC2000915 DOI: 10.1186/1471-2105-8-312
Source DB: PubMed Journal: BMC Bioinformatics ISSN: 1471-2105 Impact factor: 3.169
Changes in alignment dimensions caused by MaxAlign
| Alignment area | 1494 | 16000 |
| Number of gap-free columns | 19.4 | 162.6 |
| Number of sequences | 139.8 | 112.0 |
Average values for alignment dimensions before and after processing with MaxAlign. Estimated from 5242 Pfam alignments with between 30 and 500 sequences.
Figure 1Accuracy in phylogenetic inference. Comparison of phylogenetic accuracy obtained with different data sets. Accuracy is measured as tree similarity between the true tree (used for simulating the data set) and the reconstructed tree. Each line shows the distribution of the accuracy results from 1000 different data sets, in the form of a box plot. The box has lines at the lower quartile, median and upper quartile. The whiskers extend from each quartile to the most extreme values within 1.5 times the interquartile range. Outliers falling outside this range are marked with dots. The datasets are in the same order (from top to bottom) as in table 2: The top two rows show the original dataset without and with removal of gapped columns, respectively. The third and fourth rows show the equivalent MaxAlign datasets. The trees in the top four rows are being evaluated on the subset of sequences shared by all data sets ("Subset"), while the lower two rows show the results for original datasets when evaluated on the full set of sequences ("All").
Effects of MaxAlign and removal of gapped columns on phylogenetic accuracy
| - | - | Subset | 85.9 (0.5) | 85.4 (0.2) | 78.7 (0.3)* |
| - | + | Subset | 21.0 (0.5) | 60.1 (0.4) | 49.1 (0.4) |
| + | - | Subset | 82.8 (0.6) | 82.3 (0.2) | 78.6 (0.3)* |
| + | + | Subset | 74.6 (0.7) | 75.9 (0.3) | 72.4 (0.3) |
| - | - | All | 56.1 (0.5) | 78.4 (0.2) | 79.7 (0.3) |
| - | + | All | 10.3 (0.2) | 52.7 (0.4) | 53.2 (0.4) |
Phylogenetic accuracy for datasets with/without removal of gapped columns, and processed/not processed by MaxAlign. Accuracy is measured as the average normalized symmetric tree similarity between the true tree (used for simulating data) and the individual inferred trees, with the standard error of the mean (in %) given in parenthesis. "Subset" refers to the set of taxa (sequences) common to the original and the MaxAligned data. "All" means all taxa in the original data set. Values marked with * are the only ones whose difference is not statistically significant.
Runtimes of the phylogenetic analysis
| - | - | 17.7 | 8.4 | 408 |
| + | - | 8.1 | 5.5 | 78 |
| + | + | 4.8 | 3.7 | 50 |
Runtimes of the phylogenetic analysis. All times measured in minutes. Both application of MaxAlign and removal of gapped columns resulted in strongly decreased runtimes.
Figure 2Example of MaxAlign processing. Example alignment, before (a) and after (b) MaxAlign. In the original unprocessed alignment (a), only the three middle columns would be included in a subsequent analysis (alignment area = 3 rows × 7 columns = 21). The first three columns have the same gap pattern. After MaxAlign processing (b) (resulting in removal of sequences A and B) only the last two columns would be excluded by having gaps (alignment area = 5 rows × 6 columns = 30).
Data on the simulated datasets
| Average sequence identity | 19% | 30% | 42% | - |
| Alignment length | 1080 | 629 | 597 | 404 |
| Sequence length | 173 | 177 | 169 | 171 |
| Original number of sequences | 32 | 33 | 46 | - |
| Average number of sequences after MaxAlign | 14.1 | 22.6 | 28.8 | - |
| Average number of indels per sequence | 66.6 | 54.3 | 48.5 | 32 |
| Average length of indels | 13.6 | 8.3 | 8.8 | 7 |
Description of the simulated alignments used for testing the accuracy of phylogenetic inference with MaxAlign and removal of gapped columns, as well as the Pfam estimates used to tune the simulation parameters.
Figure 3Tree topologies used to simulate alignments. The trees used to simulated the alignments. From 1 to 3: TF101002, TF101523 and TF105969.