Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Improved K-means clustering algorithm for exploring local protein sequence motifs representing common structural property.

Literature DB >> 16220690

Improved K-means clustering algorithm for exploring local protein sequence motifs representing common structural property.

Wei Zhong¹, Gulsah Altun, Robert Harrison, Phang C Tai, Yi Pan.

Abstract

Information about local protein sequence motifs is very important to the analysis of biologically significant conserved regions of protein sequences. These conserved regions can potentially determine the diverse conformation and activities of proteins. In this work, recurring sequence motifs of proteins are explored with an improved K-means clustering algorithm on a new dataset. The structural similarity of these recurring sequence clusters to produce sequence motifs is studied in order to evaluate the relationship between sequence motifs and their structures. To the best of our knowledge, the dataset used by our research is the most updated dataset among similar studies for sequence motifs. A new greedy initialization method for the K-means algorithm is proposed to improve traditional K-means clustering techniques. The new initialization method tries to choose suitable initial points, which are well separated and have the potential to form high-quality clusters. Our experiments indicate that the improved K-means algorithm satisfactorily increases the percentage of sequence segments belonging to clusters with high structural similarity. Careful comparison of sequence motifs obtained by the improved and traditional algorithms also suggests that the improved K-means clustering algorithm may discover some relatively weak and subtle sequence motifs, which are undetectable by the traditional K-means algorithms. Many biochemical tests reported in the literature show that these sequence motifs are biologically meaningful. Experimental results also indicate that the improved K-means algorithm generates more detailed sequence motifs representing common structures than previous research. Furthermore, these motifs are universally conserved sequence patterns across protein families, overcoming some weak points of other popular sequence motifs. The satisfactory result of the experiment suggests that this new K-means algorithm may be applied to other areas of bioinformatics research in order to explore the underlying relationships between data samples more effectively.

Mesh：

Substances：
Proteins

Year: 2005 PMID： 16220690 DOI： 10.1109/tnb.2005.853667

Source DB: PubMed Journal: IEEE Trans Nanobioscience ISSN： 1536-1241 Impact factor: 2.935

Keyword Cloud
Cited

6 in total

Improved K-means clustering algorithm for exploring local protein sequence motifs representing common structural property.

1. A special local clustering algorithm for identifying the genes associated with Alzheimer's disease.

2. CATS: A Tool for Clustering the Ensemble of Intrinsically Disordered Peptides on a Flat Energy Landscape.

3. Using graph theory to analyze biological networks.

4. 3GOLD: optimized Levenshtein distance for clustering third-generation sequencing data.

Review 5. Review of Machine Learning Methods for the Prediction and Reconstruction of Metabolic Pathways.

6. Protein local 3D structure prediction by Super Granule Support Vector Machines (Super GSVM).