Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 Attribute clustering for grouping, selection, and classification of gene expression data.

Literature DB >> 17044174

Attribute clustering for grouping, selection, and classification of gene expression data.

Wai-Ho Au¹, Keith C C Chan, Andrew K C Wong, Yang Wang.

Abstract

This paper presents an attribute clustering method which is able to group genes based on their interdependence so as to mine meaningful patterns from the gene expression data. It can be used for gene grouping, selection, and classification. The partitioning of a relational table into attribute subgroups allows a small number of attributes within or across the groups to be selected for analysis. By clustering attributes, the search dimension of a data mining algorithm is reduced. The reduction of search dimension is especially important to data mining in gene expression data because such data typically consist of a huge number of genes (attributes) and a small number of gene expression profiles (tuples). Most data mining algorithms are typically developed and optimized to scale to the number of tuples instead of the number of attributes. The situation becomes even worse when the number of attributes overwhelms the number of tuples, in which case, the likelihood of reporting patterns that are actually irrelevant due to chances becomes rather high. It is for the aforementioned reasons that gene grouping and selection are important preprocessing steps for many data mining algorithms to be effective when applied to gene expression data. This paper defines the problem of attribute clustering and introduces a methodology to solving it. Our proposed method groups interdependent attributes into clusters by optimizing a criterion function derived from an information measure that reflects the interdependence between attributes. By applying our algorithm to gene expression data, meaningful clusters of genes are discovered. The grouping of genes based on attribute interdependence within group helps to capture different aspects of gene association patterns in each group. Significant genes selected from each group then contain useful information for gene expression classification and identification. To evaluate the performance of the proposed approach, we applied it to two well-known gene expression data sets and compared our results with those obtained by other methods. Our experiments show that the proposed method is able to find the meaningful clusters of genes. By selecting a subset of genes which have high multiple-interdependence with others within clusters, significant classification information can be obtained. Thus, a small pool of selected genes can be used to build classifiers with very high classification rate. From the pool, gene expressions of different categories can be identified.

Mesh：

Year: 2005 PMID： 17044174 DOI： 10.1109/TCBB.2005.17

Source DB: PubMed Journal: IEEE/ACM Trans Comput Biol Bioinform ISSN： 1545-5963 Impact factor: 3.710

Keyword Cloud
Cited

6 in total

1. A classification framework applied to cancer gene expression profiles.

Authors: Hussein Hijazi; Christina Chan
Journal: J Healthc Eng Date: 2013 Impact factor: 2.682

2. Unsupervised fuzzy pattern discovery in gene expression data.

Authors: Gene P K Wu; Keith C C Chan; Andrew K C Wong
Journal: BMC Bioinformatics Date: 2011-07-27 Impact factor: 3.169

3. An ensemble machine learning model based on multiple filtering and supervised attribute clustering algorithm for classifying cancer samples.

Authors: Shilpi Bose; Chandra Das; Abhik Banerjee; Kuntal Ghosh; Matangini Chattopadhyay; Samiran Chattopadhyay; Aishwarya Barik
Journal: PeerJ Comput Sci Date: 2021-09-16

Attribute clustering for grouping, selection, and classification of gene expression data.

1. A classification framework applied to cancer gene expression profiles.

2. Unsupervised fuzzy pattern discovery in gene expression data.

3. An ensemble machine learning model based on multiple filtering and supervised attribute clustering algorithm for classifying cancer samples.

4. Statistical discovery of site inter-dependencies in sub-molecular hierarchical protein structuring.

5. Gene selection for cancer classification with the help of bees.

6. The ability to classify patients based on gene-expression data varies by algorithm and performance metric.