Literature DB >> 17044174

Attribute clustering for grouping, selection, and classification of gene expression data.

Wai-Ho Au1, Keith C C Chan, Andrew K C Wong, Yang Wang.   

Abstract

This paper presents an attribute clustering method which is able to group genes based on their interdependence so as to mine meaningful patterns from the gene expression data. It can be used for gene grouping, selection, and classification. The partitioning of a relational table into attribute subgroups allows a small number of attributes within or across the groups to be selected for analysis. By clustering attributes, the search dimension of a data mining algorithm is reduced. The reduction of search dimension is especially important to data mining in gene expression data because such data typically consist of a huge number of genes (attributes) and a small number of gene expression profiles (tuples). Most data mining algorithms are typically developed and optimized to scale to the number of tuples instead of the number of attributes. The situation becomes even worse when the number of attributes overwhelms the number of tuples, in which case, the likelihood of reporting patterns that are actually irrelevant due to chances becomes rather high. It is for the aforementioned reasons that gene grouping and selection are important preprocessing steps for many data mining algorithms to be effective when applied to gene expression data. This paper defines the problem of attribute clustering and introduces a methodology to solving it. Our proposed method groups interdependent attributes into clusters by optimizing a criterion function derived from an information measure that reflects the interdependence between attributes. By applying our algorithm to gene expression data, meaningful clusters of genes are discovered. The grouping of genes based on attribute interdependence within group helps to capture different aspects of gene association patterns in each group. Significant genes selected from each group then contain useful information for gene expression classification and identification. To evaluate the performance of the proposed approach, we applied it to two well-known gene expression data sets and compared our results with those obtained by other methods. Our experiments show that the proposed method is able to find the meaningful clusters of genes. By selecting a subset of genes which have high multiple-interdependence with others within clusters, significant classification information can be obtained. Thus, a small pool of selected genes can be used to build classifiers with very high classification rate. From the pool, gene expressions of different categories can be identified.

Mesh:

Year:  2005        PMID: 17044174     DOI: 10.1109/TCBB.2005.17

Source DB:  PubMed          Journal:  IEEE/ACM Trans Comput Biol Bioinform        ISSN: 1545-5963            Impact factor:   3.710


  6 in total

1.  A classification framework applied to cancer gene expression profiles.

Authors:  Hussein Hijazi; Christina Chan
Journal:  J Healthc Eng       Date:  2013       Impact factor: 2.682

2.  Unsupervised fuzzy pattern discovery in gene expression data.

Authors:  Gene P K Wu; Keith C C Chan; Andrew K C Wong
Journal:  BMC Bioinformatics       Date:  2011-07-27       Impact factor: 3.169

3.  An ensemble machine learning model based on multiple filtering and supervised attribute clustering algorithm for classifying cancer samples.

Authors:  Shilpi Bose; Chandra Das; Abhik Banerjee; Kuntal Ghosh; Matangini Chattopadhyay; Samiran Chattopadhyay; Aishwarya Barik
Journal:  PeerJ Comput Sci       Date:  2021-09-16

4.  Statistical discovery of site inter-dependencies in sub-molecular hierarchical protein structuring.

Authors:  Kirk K Durston; David Ky Chiu; Andrew Kc Wong; Gary Cl Li
Journal:  EURASIP J Bioinform Syst Biol       Date:  2012-07-13

5.  Gene selection for cancer classification with the help of bees.

Authors:  Johra Muhammad Moosa; Rameen Shakur; Mohammad Kaykobad; Mohammad Sohel Rahman
Journal:  BMC Med Genomics       Date:  2016-08-10       Impact factor: 3.063

6.  The ability to classify patients based on gene-expression data varies by algorithm and performance metric.

Authors:  Stephen R Piccolo; Avery Mecham; Nathan P Golightly; Jérémie L Johnson; Dustin B Miller
Journal:  PLoS Comput Biol       Date:  2022-03-11       Impact factor: 4.475

  6 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.