| Literature DB >> 16524483 |
Haiying Wang1, Huiru Zheng, David Simpson, Francisco Azuaje.
Abstract
BACKGROUND: Retinal photoreceptors are highly specialised cells, which detect light and are central to mammalian vision. Many retinal diseases occur as a result of inherited dysfunction of the rod and cone photoreceptor cells. Development and maintenance of photoreceptors requires appropriate regulation of the many genes specifically or highly expressed in these cells. Over the last decades, different experimental approaches have been developed to identify photoreceptor enriched genes. Recent progress in RNA analysis technology has generated large amounts of gene expression data relevant to retinal development. This paper assesses a machine learning methodology for supporting the identification of photoreceptor enriched genes based on expression data.Entities:
Mesh:
Substances:
Year: 2006 PMID: 16524483 PMCID: PMC1421439 DOI: 10.1186/1471-2105-7-116
Source DB: PubMed Journal: BMC Bioinformatics ISSN: 1471-2105 Impact factor: 3.169
Figure 1A hierarchical tree for 14 SAGE libraries based on analysis of 1118 tags. The hierarchical tree was generated based on 1118 tags highly expressed in the ONL library using Pearson correlation-based hierarchical clustering method. Five clusters are obtained when cutting the dendrogram at level A.
Figure 2A hierarchical tree for 14 SAGE libraries based on analysis of 261 PR-enriched tags. This hierarchical tree provides further insights. The split at birth is less marked. The P10.5 Crx-/- library is clustered with the P6.5 library rather than with its wild type counterpart.
Figure 3The mean values of 5-runs adjusted FOM calculations against the number of clusters. An open-source implementation of FOM provided by the Institute for Genomic Research (TIGR) [22] was used to calculate the adjusted FOM for k-means algorithm with Euclidean distance on normalised SAGE data. For a given tag, the abundance in each SAGE library was rescaled to make the sum of tag counts across all 14 libraries equal to one.
Figure 4Poisson model-based clustering analysis for photoreceptor gene expression using 10 clusters. SAGE libraries are plotted on the x-axis. Numbers one to fourteen represent the fourteen SAGE libraries mentioned in Method Section, i.e. 1: hypo; 2: 3t3; 3: E12.5; 4: 14.5; 5: E16.5; 6: E18.5; 7: P0.5; 8: P2.5; 9: P4.5; 10: P6.5; 11: P10.5Crx-/-; 12: P10.5Crx+/+; 13: Adult; 14: ONL. Tag abundance is shown on the y-axis. Data were normalized before plotting. Each tag from the 14 libraries was rescaled to make the sum of the expression values equal to one. Different colors represent different tags.
Summary of Poisson-based analysis for the SAGE data.
| Cluster | No. of Tags | Description of cluster profile | No. of Tags validated | No. of Tags validated by Blackshaw | ||
| PR-enriched | non-PR-enriched | PR-enriched | non-PR-enriched | |||
| 1 | 105 | Varying expression | 0 | 0 | 5 | 1 |
| 2 | 196 | Increasing during postnatal development | 26 | 1 | 114 | 13 |
| 3 | 53 | Peak expression occurred during embryonic development | 0 | 1 | 1 | 6 |
| 4 | 280 | Flat | 1 | 0 | 19 | 11 |
| 5 | 63 | Peak expression occurred during embryonic development | 0 | 0 | 0 | 4 |
| 6 | 102 | Increasingly sharp during late postnatal development | 34 | 0 | 20 | 4 |
| 7 | 27 | Peak expression in adult hypothalamus | 0 | 0 | 0 | 0 |
| 8 | 55 | Peak expression occurred around P4.5 and ONL | 1 | 0 | 4 | 2 |
| 9 | 132 | Peak expression in NIH-3T3 fibroblast cells | 1 | 8 | 16 | 5 |
| 10 | 132 | Varying expression | 1 | 0 | 18 | 7 |
A list of rules extracted from 324 tags using the Apriori algorithm.
| Enriched in PR? <= criterion 4 (234:72.2%, 0.829) |
| Enriched in PR? <= criterion 1 (218:67.3%, 0.899) |
| Enriched in PR? <= criterion 3 (177:54.6%, 0.842) |
| Enriched in PR? <= criterion 2 (175:54.0%, 0.937) |
| Enriched in PR? <= criterion 4 & criterion 1 (165:50.9%, 0.909) |
| Enriched in PR? <= criterion 4 & criterion 3 (129:39.8%, 0.884) |
| Enriched in PR? <= criterion 4 & criterion 2 (117:36.1%, 0.949) |
| Enriched in PR? <= criterion 1 & criterion 3 (126:38.9%, 0.929) |
| Enriched in PR? <= criterion 1 & criterion 2 (143:44.1%, 0.944) |
| Enriched in PR? <= criterion 3 & criterion 2 (105:32.4%, 0.943) |
| Enriched in PR? <= criterion 4 & criterion 1 & criterion 3 (92:28.4%, 0.946) |
| Enriched in PR? <= criterion 4 & criterion 1 & criterion 2 (100:30.9%, 0.96) |
| Enriched in PR? <= criterion 4 & criterion 3 & criterion 2 (70:21.6%, 0. |
| Enriched in PR? <= criterion 1 & criterion 3 & criterion 2 (90:27.8%, 0.944) |
| Enriched in PR?<= criterion 4 & criterion 1 & criterion 3 & criterion 2 (59:18.2%, 0.966) |
Each rule is shown in the following format: Enriched in PR? <= criterion 1 & criterion 2 & ... &criterion n where the rule is interpreted as "for tags that meet criterion 1 through criterion n, they are likely to be PR-enriched genes." The numbers shown at the end of each rule indicate the number of tags to which the rule applies (Support) and the proportion of those tags for which the rule is true (Confidence). Support is reported both as number of tags and percentage of total tags, separated by a colon.
Prediction results of 10-fold cross validation for three classifiers using random over-sampling method. The total number of SAGE tags analyzed is 522, in which 261 are PR-enriched. Each tag is represented by 14 SAGE libraries.
| Method | PR-enriched | non-PR-enriched | |||
| KStar | 91.2 | 99.5 | 82.8 | 99.6 | 85.2 |
| C4.5 | 91.0 | 97.3 | 84.3 | 97.7 | 86.1 |
| MLP | 66.5 | 61.6 | 87.4 | 45.6 | 78.3 |
Prediction results of 10-fold cross validation for three classifiers using random under-sampling method. The total number of SAGE tags analyzed is 126, in which 63 are PR enriched. Each tag is represented by 14 SAGE libraries.
| Method | PR-enriched | non-PR-enriched | |||
| Kstar | 61.9 | 61.5 | 63.5 | 60.3 | 62.3 |
| C4.5 | 66.7 | 66.7 | 66.7 | 66.7 | 66.7 |
| MLP | 65.1 | 60.9 | 84.1 | 46.0 | 74.4 |
Figure 5The ROC curves for three classifiers using random over-sampling and under-sampling methods. Figure (a1) KStar with over-sampling method; (a2) KStar with under-sampling method; (b1) C4.5 with over-sampling method; (b2) C4.5 with under-sampling method; (c1) MLP with over-sampling method; (c2) MLP with under-sampling method. The blue and yellow lines in ROC curves represent the different threshold values which are used to generate ROC.
The effect of class distribution on the performance of the classifier.
| Class distribution (PR-enriched : non-PR-enriched) | PR enriched | non-PR enriched | |||
| 1:1 | 91.2 | 99.5 | 82.8 | 99.6 | 85.2 |
| 2:1 | 89.8 | 95.7 | 85.8 | 94.9 | 83.6 |
| 3:1 | 89.1 | 94.3 | 88.9 | 89.6 | 80.7 |
| 4:1 | 84.1 | 88.8 | 91.2 | 58.3 | 64.6 |
| Original (261:63) | 81.2 | 86.2 | 91.2 | 39.7 | 52.1 |
The class distribution is obtained based on random over sampling method. A 10-fold cross validation was carried out to estimate the true classification error. Class distribution is represented as the number of PR-enriched tags against the number of non-PR-enriched tags. There are 261 PR-enriched and 63 non-PR-enriched tags in the original dataset. In all resampled datasets, the number of PR-enriched tags is as also equal to 261.
Distribution of tags under study in terms of functional classification
| Class | Number of tags | ||
| PR-enriched | TRUE | 197 | 261 |
| KNOWN | 64 | ||
| non-PR-enriched | FALSE | 53 | 63 |
| KNOWN FALSE | 10 | ||
| N.D. | 708 | ||
| UNKNOWN | 86 | ||
| Total | 1118 | ||
Under the Class column, TRUE and FALSE stand for genes validated by Blackshaw et al. using ISH; KNOWN and KNOWN FALSE represent genes validated prior to Blackshaw et al. 's study; N.D. stands for tags not validated in Blackshaw et al.'s study; and UNKNOWN includes tags which did not correspond to any identifiable transcript [7].