| Literature DB >> 15271224 |
Dariusz Plewczynski1, Leszek Rychlewski, Yuzhen Ye, Lukasz Jaroszewski, Adam Godzik.
Abstract
BACKGROUND: Defining blocks forming the global protein structure on the basis of local structural regularity is a very fruitful idea, extensively used in description, and prediction of structure from only sequence information. Over many years the secondary structure elements were used as available building blocks with great success. Specially prepared sets of possible structural motifs can be used to describe similarity between very distant, non-homologous proteins. The reason for utilizing the structural information in the description of proteins is straightforward. Structural comparison is able to detect approximately twice as many distant relationships as sequence comparison at the same error rate.Entities:
Mesh:
Substances:
Year: 2004 PMID: 15271224 PMCID: PMC497040 DOI: 10.1186/1471-2105-5-98
Source DB: PubMed Journal: BMC Bioinformatics ISSN: 1471-2105 Impact factor: 3.169
Figure 1The FRAGlib fragments database is build from ASTRAL representative subset of SCOP database using 40% sequence similarity threshold (see right picture). We store the symbolized local structure representation codes of each fragment together with the homology sequence profile (see left picture). Both are dissected from the SLSR codes and homology profile of a parent protein. The string of SLSR codes representing the local structure of the Cα chain in the phi-phi space. We remove from the fragments database all identical in terms of both SLSR codes and sequence homology profile fragments. On the left picture we present the FRAGlib module for prediction of local structural segments using homology profile similarity and the fragments database. The Query protein is dissected into short parts (from 7 up to 19 residues long). For each part the similarity search is performed. Any member of the fragments database which is similar in terms of homology sequence profile similarity is added to the list of predicted structures for this short part of query protein. This list is then sorted and cut after arbitrary chosen 20th position. If the highest score of predicted fragments is below the user's cut-off value whole prediction is discarder. In the end some of parts of a query protein are covered by list of 20 fragments from the database. They are called the predicted local structural segments (PLSSs).
Figure 2We present here the flowchart of SEA/FRAGlib integrated Web service. The server is based on two modules: the FRAGlib prediction of LSSs and the SEA algorithm for building an alignment between two proteins using comparison of two networks of predicted segments for both of them.
General performance of classical methods for building alignments together with segment alignment algorithm incorporating different local structure diversities.
| CE | SEAT | SEAI | SEAloc | BLAST | ALIGN | FFAS | ||||
| Family (409 pairs) | shift | average | 0.61 | 0.56 | 0.49 | 0.44 | 0.48 | 0.49 | ||
| >0.9 | 73 | 69 | 47 | 51 | 60 | 43 | ||||
| >0.7 | 207 | 199 | 152 | 146 | 165 | 161 | ||||
| >0.5 | 282 | 260 | 215 | 197 | 228 | 227 | ||||
| RMS | ≤ 3.0 | 257 | 95 | 82 | 63 | 77 | 54 | 40 | ||
| ≤ 5.0 | 397 | 237 | 184 | 147 | 157 | 138 | 118 | |||
| ≤ 8.0 | 408 | 294 | 248 | 231 | 196 | 206 | 194 | |||
| all | 409 | 345 | 404 | 366 | 232 | 372 | 409 | |||
| len | 1 | 0.84 | 1.08 | 0.87 | 0.56 | 0.99 | 1.18 |
Family-level benchmark for SEA algorithm using FRAGlib's prediction of LSSs (SEAF) is compared with SEAI (SEA algorithm using I-sites library), SEAT, SEAloc (local single predicted structures), and other classical tools: CE, BLAST, ALIGN and FFAS. The 'average' is the shift score averaged over all the alignments of the whole subset. The numbers of protein pairs with a shift score or RMSD larger than a certain cut-off value in the subset are listed in columns for each program. The counting based on RMSD requires the length of the alignment to be longer than half of its corresponding structural alignment. The 'all' stands for all the alignments with alignment length no shorter than half of the structural alignments. We use the CE for building reference alignments for shift score calculation, as an example of purely structural alignment tool. The 'len' stands for the average alignment length (predicted aligned position / aligned position in reference alignment from CE). We can see that our method provides very long alignments with relatively good overall score. The difference in the values between SEAT and SEAF is explained by different lengths of these alignments.