| Literature DB >> 19038021 |
Marcilio C P de Souto1, Ivan G Costa, Daniel S A de Araujo, Teresa B Ludermir, Alexander Schliep.
Abstract
BACKGROUND: The use of clustering methods for the discovery of <al">span class="Disease">cancer subtypes has drawn a great deal of attention in the scientific community. While bioinformaticians have proposed new clustering methods that take advantage of characteristics of the gene expression data, the medical community has a preference for using "classic" clustering methods. There have been no studies thus far performing a large-scale evaluation of different clustering methods in this context. RESULTS/Entities:
Mesh:
Substances:
Year: 2008 PMID: 19038021 PMCID: PMC2632677 DOI: 10.1186/1471-2105-9-497
Source DB: PubMed Journal: BMC Bioinformatics ISSN: 1471-2105 Impact factor: 3.169
Data set description
| Armstrong-V1 [ | Affy | Blood | 72 | 2 | 24,48 | 12582 | 1081 |
| Armstrong-V2 [ | Affy | Blood | 72 | 3 | 24,20,28 | 12582 | 2194 |
| Bhattacharjee [ | Affy | Lung | 203 | 5 | 139,17,6,21,20 | 12600 | 1543 |
| Chowdary [ | Affy | Breast, Colon | 104 | 2 | 62,42 | 22283 | 182 |
| Dyrskjot [ | Affy | Bladder | 40 | 3 | 9,20,11 | 7129 | 1203 |
| Golub-V1 [ | Affy | Bone marrow | 72 | 2 | 47,25 | 7129 | 1877 |
| Golub-V2 [ | Affy | Bone marrow | 72 | 3 | 38,9,25 | 7129 | 1877 |
| Gordon [ | Affy | Lung | 181 | 2 | 31,150 | 12533 | 1626 |
| Laiho [ | Affy | Colon | 37 | 2 | 8,29 | 22883 | 2202 |
| Nutt-V1 [ | Affy | Brain | 50 | 4 | 14,7,14,15 | 12625 | 1377 |
| Nutt-V2 [ | Affy | Brain | 28 | 2 | 14,14 | 12625 | 1070 |
| Nutt-V3 [ | Affy | Brain | 22 | 2 | 7,15 | 12625 | 1152 |
| Pomeroy-V1 [ | Affy | Brain | 34 | 2 | 25,9 | 7129 | 857 |
| Pomeroy-V2 [ | Affy | Brain | 42 | 5 | 10,10,10,4,8 | 7129 | 1379 |
| Ramaswamy [ | Affy | Multi-tissue | 190 | 14 | 11,10,11,11,22,10,11,10,30,11,11,11,11,20 | 16063 | 1363 |
| Shipp [ | Affy | Blood | 77 | 2 | 58,19 | 7129 | 798 |
| Singh [ | Affy | Prostate | 102 | 2 | 58,19 | 12600 | 339 |
| Su [ | Affy | Multi-tissue | 174 | 10 | 26,8,26,23,12,11,7,27,6,28 | 12533 | 1571 |
| West [ | Affy | Breast | 49 | 2 | 25,24 | 7129 | 1198 |
| Yeoh-V1 [ | Affy | Bone marrow | 248 | 2 | 43,205 | 12625 | 2526 |
| Yeoh-V2 [ | Affy | Bone marrow | 248 | 6 | 15,27,64,20,79,43 | 12625 | 2526 |
| Alizadeh-V1 [ | cDNA | Blood | 42 | 2 | 21,21 | 4022 | 1095 |
| Alizadeh-V2 [ | cDNA | Blood | 62 | 3 | 42,9,11 | 4022 | 2093 |
| Alizadeh-V3 [ | cDNA | Blood | 62 | 4 | 21,21,9,11 | 4022 | 2093 |
| Bittner [ | cDNA | Skin | 38 | 2 | 19, 19 | 8067 | 2201 |
| Bredel [ | cDNA | Brain | 50 | 3 | 31,14,5 | 41472 | 1739 |
| Chen [ | cDNA | Liver | 180 | 2 | 104,76 | 22699 | 85 |
| Garber [ | cDNA | Lung | 66 | 4 | 17,40,4,5 | 24192 | 4553 |
| Khan [ | cDNA | Multi-tissue | 83 | 4 | 29,11,18,25 | 6567 | 1069 |
| Lapointe-V1 [ | cDNA | Prostate | 69 | 3 | 11,39,19 | 42640 | 1625 |
| Lapoint-V2 [ | cDNA | Prostate | 110 | 4 | 11,39,19,41 | 42640 | 2496 |
| Liang [ | cDNA | Brain | 37 | 3 | 28,6,3 | 24192 | 1411 |
| Risinger [ | cDNA | Endometrium | 42 | 4 | 13,3,19,7 | 8872 | 1771 |
| Tomlins-V1 [ | cDNA | Prostate | 104 | 5 | 27,20,32,13,12 | 20000 | 2315 |
| Tomlins-V2 [ | cDNA | Prostate | 92 | 4 | 27,20,32,13 | 20000 | 1288 |
The data sets present different values for features such as type of microarray chip (second column), tissue type (third column), number of samples (fourth column), number of classes (fifth column), distribution of samples within the classes (sixth column), dimensionality (seventh column) and dimensionality after feature selection (last column).
Figure 1Affymetrix data sets: mean of the cR. We display the mean of the cR for the partitions with the number of clusters equal to the actual number of classes (a) and the mean of the best cR found (b) for Affymetrix data sets. Missing bars correspond to combinations of methods and proximity measures not evaluated.
Figure 2cDNA data sets: mean of the cR. We display the mean of the cR for the partitions with the number of clusters equal to the actual number of classes (a) and the mean of the best cR found (b) for cDNA data sets. Missing bars correspond to combinations of methods and proximity measures not evaluated.
Figure 3Affymetrix data sets: difference between the actual number of classes and the number of clusters in the partition solutions with best cR. We display the mean of the difference between the actual number of classes and the number of clusters for the best partition found for Affymetrix data sets. Missing bars correspond to combinations of methods and proximity measures not evaluated.
Figure 4cDNA data sets: difference between the actual number of classes and the number of clusters in the partition solutions with best cR. We display the mean of the difference between the actual number of classes and the number of clusters for the best partition found for cDNA data sets. Missing bars correspond to combinations of methods and proximity measures not evaluated.
Figure 5Hierarchical Clustering and . We display the red and green plots for (a) the hierarchical clustering and (b) the k-means for the data set Alizadeh-V2. Columns correspond to genes and rows correspond to cancer samples. The samples are labeled according to one of three classes: diffuse large B-cell lymphoma (DLBCL), follicular lymphoma (FL) and chronic lymphocytic leukemia (CLL). In the case of k-means, the number of clusters was set at three. Likewise, for hierarchical clustering, the tree was cut so as to return three clusters, corresponding to the red, light blue and black sub-trees.
Figure 6PCA plot for Alizadeh-V2. We display a scatter plot with the two first largest components of a PCA for Alizadeh-V2. Colors indicate the three classes in the data: diffuse large B-cell lymphoma in red (DLBCL), follicular lymphoma in green (FL) and chronic lymphocytic leukemia in blue(CLL).