| Literature DB >> 22125455 |
Li-Ping Tian1, Li-Zhi Liu, Qian-Wei Zhang, Fang-Xiang Wu.
Abstract
Clustering periodically expressed genes from their time-course expression data could help understand the molecular mechanism of those biological processes. In this paper, we propose a nonlinear model-based clustering method for periodically expressed gene profiles. As periodically expressed genes are associated with periodic biological processes, the proposed method naturally assumes that a periodically expressed gene dataset is generated by a number of periodical processes. Each periodical process is modelled by a linear combination of trigonometric sine and cosine functions in time plus a Gaussian noise term. A two stage method is proposed to estimate the model parameter, and a relocation-iteration algorithm is employed to assign each gene to an appropriate cluster. A bootstrapping method and an average adjusted Rand index (AARI) are employed to measure the quality of clustering. One synthetic dataset and two biological datasets were employed to evaluate the performance of the proposed method. The results show that our method allows the better quality clustering than other clustering methods (e.g., k-means) for periodically expressed gene data, and thus it is an effective cluster analysis method for periodically expressed gene data.Entities:
Keywords: Gene expression data; average adjusted Rand index; clustering; nonlinear model; periodicall expressed genes
Mesh:
Year: 2011 PMID: 22125455 PMCID: PMC3217600 DOI: 10.1100/2011/520498
Source DB: PubMed Journal: ScientificWorldJournal ISSN: 1537-744X
Algorithm 1Algorithm for nonlinear model-based clustering.
Contingency table for two partitions of n objects.
|
|
|
| Total | ||
|---|---|---|---|---|---|
|
|
|
| ⋯ |
|
|
|
|
|
| ⋯ |
|
|
| ⋮ | ⋮ | ⋮ | ⋮ | ⋮ | |
|
|
|
| ⋯ |
|
|
|
| |||||
| Total |
|
| ⋯ |
|
|
Algorithm 2The procedure for evaluating proposed clustering method.
Parameters for synthetic data.
| Cluster 1 | Cluster 2 | Cluster 3 | Cluster 4 | Cluster 5 | |
|---|---|---|---|---|---|
|
| 3.4397 | 6.9227 | 9.9126 | 12.1470 | 14.8819 |
|
| 2.3705 | 6.0603 | 8.9280 | 12.2379 | 14.8195 |
|
| 5.1516 | 3.1085 | 1.9359 | 1.5344 | 1.2413 |
|
| 0.4000 | 0.4000 | 0.4000 | 0.4000 | 0.4000 |
|
| 136 | 300 | 279 | 239 | 120 |
The values of AARI for different clustering methods on synthetic data.
| No. of clusters | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|
| Random initial | 0.2638 | 0.4391 | 0.7734 | 0.9778 | 0.8766 | 0.8361 | 0.8140 | 0.7991 | 0.7895 |
|
| 0.2586 | 0.5088 | 0.6335 | 0.6970 | 0.7794 | 0.7731 | 0.7620 | 0.6758 | 0.6904 |
|
| 0.2310 | 0.4391 | 0.7889 | 0.9913 | 0.8824 | 0.8328 | 0.8185 | 0.8027 | 0.7771 |
Figure 1Plot of ARRI over different number of clusters for synthetic dataset.
The number of genes after filtering.
|
| 0.10 | 0.20 |
|---|---|---|
| ELU | 691 | 1207 |
| BAC | 471 | 658 |
Figure 2Plot of ARRI over different number of clusters for ELU after filtering.
Figure 3Plot of ARRI over different number of clusters for BAC after filtering.