Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 FastaHerder2: Four Ways to Research Protein Function and Evolution with Clustering and Clustered Databases.

Literature DB >> 26828375

FastaHerder2: Four Ways to Research Protein Function and Evolution with Clustering and Clustered Databases.

Pablo Mier^1,2, Miguel A Andrade-Navarro^1,2.

Abstract

The accelerated growth of protein databases offers great possibilities for the study of protein function using sequence similarity and conservation. However, the huge number of sequences deposited in these databases requires new ways of analyzing and organizing the data. It is necessary to group the many very similar sequences, creating clusters with automated derived annotations useful to understand their function, evolution, and level of experimental evidence. We developed an algorithm called FastaHerder2, which can cluster any protein database, putting together very similar protein sequences based on near-full-length similarity and/or high threshold of sequence identity. We compressed 50 reference proteomes, along with the SwissProt database, which we could compress by 74.7%. The clustering algorithm was benchmarked using OrthoBench and compared with FASTA HERDER, a previous version of the algorithm, showing that FastaHerder2 can cluster a set of proteins yielding a high compression, with a lower error rate than its predecessor. We illustrate the use of FastaHerder2 to detect biologically relevant functional features in protein families. With our approach we seek to promote a modern view and usage of the protein sequence databases more appropriate to the postgenomic era.

Keywords: cluster analysis; clustering; computational biology; data mining; databases

Mesh：

Substances：
Proteome

Year: 2016 PMID： 26828375 DOI： 10.1089/cmb.2015.0191

Source DB: PubMed Journal: J Comput Biol ISSN： 1066-5277 Impact factor: 1.479

Keyword Cloud
Cited

8 in total

FastaHerder2: Four Ways to Research Protein Function and Evolution with Clustering and Clustered Databases.

1. Glutamine Codon Usage and polyQ Evolution in Primates Depend on the Q Stretch Length.

2. dAPE: a web server to detect homorepeats and follow their evolution.

3. The Protein Structure Context of PolyQ Regions.

4. Geometric characterisation of disease modules.

5. Manifold learning and maximum likelihood estimation for hyperbolic network embedding.

6. Efficient embedding of complex networks to hyperbolic space via their Laplacian.

7. CABRA: Cluster and Annotate Blast Results Algorithm.

8. The latent geometry of the human protein interaction network.