Warning: Undefined array key "mm" in /www/wwwroot/www.ai-bt.com/si.php on line 10 Deprecated: trim(): Passing null to parameter #1 ($string) of type string is deprecated in /www/wwwroot/www.ai-bt.com/si.php on line 10 How well do we understand the clusters found in microarray data?

Literature DB >> 12611631

How well do we understand the clusters found in microarray data?

Abstract

We wished to quantify the state-of-the-art of our understanding of clusters in microarray data. To do this we systematically compared the clusters produced on sets of microarray data using a representative set of clustering algorithms (hierarchical, k-means, and a modified version of QT_CLUST) with the annotation schemes MIPS, GeneOntology and GenProtEC. We assumed that if a cluster reflected known biology its members would share related ontological annotations. This assumption is the basis of "guilt-by-association" and is commonly used to assign the putative function of proteins. To statistically measure the relationship between cluster and annotation we developed a new predictive discriminatory measure. We found that the clusters found in microarray data do not in general agree with functional annotation classes. Although many statistically significant relationships can be found, the majority of clusters are not related to known biology (as described in annotation ontologies). This implies that use of guilt-by-association is not supported by annotation ontologies. Depending on the estimate of the amount of noise in the data, our results suggest that bioinformatics has only codified a small proportion of the biological knowledge required to understand microarray data.

Mesh：

Substances：
Fungal Proteins
Proteome

Year: 2002 PMID： 12611631

Source DB: PubMed Journal: In Silico Biol ISSN： 1386-6338

Keyword Cloud
Cited

8 in total

1. The FunCat, a functional annotation scheme for systematic classification of proteins from whole genomes.

Authors: Andreas Ruepp; Alfred Zollner; Dieter Maier; Kaj Albermann; Jean Hani; Martin Mokrejs; Igor Tetko; Ulrich Güldener; Gertrud Mannhaupt; Martin Münsterkötter; H Werner Mewes
Journal: Nucleic Acids Res Date: 2004-10-14 Impact factor: 16.971

Review 2. Gene expression profiling and the use of genome-scale in silico models of Escherichia coli for analysis: providing context for content.

Authors: Nathan E Lewis; Byung-Kwan Cho; Eric M Knight; Bernhard O Palsson
Journal: J Bacteriol Date: 2009-04-10 Impact factor: 3.490

3. Systematic survey reveals general applicability of "guilt-by-association" within gene coexpression networks.

Authors: Cecily J Wolfe; Isaac S Kohane; Atul J Butte
Journal: BMC Bioinformatics Date: 2005-09-14 Impact factor: 3.169

4. The Mouse Functional Genome Database (MfunGD): functional annotation of proteins in the light of their cellular context.

Authors: Andreas Ruepp; Octave Noubibou Doudieu; Jos van den Oever; Barbara Brauner; Irmtraud Dunger-Kaltenbach; Gisela Fobo; Goar Frishman; Corinna Montrone; Christine Skornia; Steffi Wanka; Thomas Rattei; Philipp Pagel; Louise Riley; Dmitrij Frishman; Dimitrij Surmeli; Igor V Tetko; Matthias Oesterheld; Volker Stümpflen; H Werner Mewes
Journal: Nucleic Acids Res Date: 2006-01-01 Impact factor: 16.971

How well do we understand the clusters found in microarray data?

1. The FunCat, a functional annotation scheme for systematic classification of proteins from whole genomes.

Review 2. Gene expression profiling and the use of genome-scale in silico models of Escherichia coli for analysis: providing context for content.

3. Systematic survey reveals general applicability of "guilt-by-association" within gene coexpression networks.

4. The Mouse Functional Genome Database (MfunGD): functional annotation of proteins in the light of their cellular context.

5. Integrated biclustering of heterogeneous genome-wide datasets for the inference of global regulatory networks.

6. Large-scale clustering of CAGE tag expression data.

7. Recursive cluster elimination (RCE) for classification and feature selection from gene expression data.

8. Public databases and software for the pathway analysis of cancer genomes.