| Literature DB >> 24260172 |
Adi L Tarca1, Gaurav Bhatti, Roberto Romero.
Abstract
Identification of functional sets of genes associated with conditions of interest from omics data was first reported in 1999, and since, a plethora of enrichment methods were published for systematic analysis of gene sets collections including Gene Ontology and biological pathways. Despite their widespread usage in reducing the complexity of omics experiment results, their performance is poorly understood. Leveraging the existence of disease specific gene sets in KEGG and Metacore® databases, we compared the performance of sixteen methods under relaxed assumptions while using 42 real datasets (over 1,400 samples). Most of the methods ranked high the gene sets designed for specific diseases whenever samples from affected individuals were compared against controls via microarrays. The top methods for gene set prioritization were different from the top ones in terms of sensitivity, and four of the sixteen methods had large false positives rates assessed by permuting the phenotype of the samples. The best overall methods among those that generated reasonably low false positive rates, when permuting phenotypes, were PLAGE, GLOBALTEST, and PADOG. The best method in the category that generated higher than expected false positives was MRGSE.Entities:
Mesh:
Year: 2013 PMID: 24260172 PMCID: PMC3829842 DOI: 10.1371/journal.pone.0079217
Source DB: PubMed Journal: PLoS One ISSN: 1932-6203 Impact factor: 3.240
Figure 1Procedure used to compare 16 gene set analysis methods.
42 microarray datasets were used, each studying a phenotype that has a corresponding KEGG or Metacore disease pathway, that we call target pathway. Each method was applied on each datasets and the p-value and rank of the target pathway in each dataset was used to compare the methods.
Figure 2A comparison of sensitivity and prioritization ability of 16 gene set analysis methods.
Each box contains 42 data points representing the p-value (left) and the rank (%) (right) that the target pathway received from a given method when using as input an independent dataset and a collection of gene sets (either KEGG or Metacore). Since the target pathways were designed by KEGG and Metacore for those diseases we expected that, in average, they will be found relevant by the different methods. Methods are ranked from best to worst according to the median p-value (left) and median rank (right).
Figure 3False positive rates produced by 16 gene sets analysis methods.
The null hypothesis is simulated by analyzing phenotype permuted versions of the original datasets. The percentage of all pathways found significant at different significance levels (α) is reported for each method with a vertical bar. The horizontal lines denote the expected levels of false positives at each α level. Note the logarithmic scale.
Ranking of gene set analysis methods.
| Raw Data | Z-scores | Overall methodrank in category | ||||||||
| Method | Sampling type | Category | Sensitivity | Prioritization | Specificity | Sensitivity | Prioritization | Specificity | Sum Z-scores | |
| med. p | med. rank(%) | FP (α = 1%) | med. p | med. rank(%) | FP (α = 1%) | |||||
|
| subject | I | 0.0022 | 25.0 | 1.1% | −1.5 | −0.4 | Not used | −1.86 | 1 |
|
| subject | I | 0.0001 | 27.9 | 2.0% | −1.5 | −0.2 | Not used | −1.69 | 2 |
|
| subject | I | 0.0960 | 9.7 | 2.5% | 0.0 | −1.5 | Not used | −1.45 | 3 |
|
| gene | I | 0.0732 | 18.3 | 2.5% | −0.4 | −0.9 | Not used | −1.21 | 4 |
|
| subject | I | 0.1065 | 18.8 | 1.3% | 0.2 | −0.8 | Not used | −0.64 | 5 |
|
| subject | I | 0.0565 | 38.0 | 0.9% | −0.6 | −0.5 | Not used | −0.09 | 6 |
|
| subject | I | 0.1420 | 21.0 | 1.3% | 0.7 | −0.7 | Not used | 0.07 | 7 |
|
| subject | I | 0.0808 | 40.3 | 1.0% | −0.2 | 0.7 | Not used | 0.45 | 8 |
|
| subject | I | 0.0950 | 39.8 | 1.0% | 0.0 | 0.7 | Not used | 0.65 | 9 |
|
| subject | I | 0.1801 | 33.1 | 2.3% | 1.3 | 0.2 | Not used | 1.52 | 10 |
|
| subject | I | 0.1986 | 51.5 | 1.1% | 1.6 | 1.5 | Not used | 3.10 | 11 |
|
| subject | I | 0.3126 | 43.0 | 0.5% | 3.4 | 0.9 | Not used | 4.30 | 12 |
|
| gene | II | 0.0100 | 18.8 | 4.9% | −0.59 | −1.68 | −1.27 | −3.54 | 1 |
|
| gene | II | 0.0644 | 36.2 | 15.8% | 0.59 | 0.02 | −0.08 | 0.53 | 2 |
|
| gene | II | 0.0024 | 35.9 | 37.9% | −0.76 | −0.02 | 2.33 | 1.56 | 3 |
|
| gene | II | 0.1165 | 49.7 | 17.2% | 1.72 | 1.33 | 0.08 | 3.14 | 4 |
Surrogate sensitivity, prioritization ability and specificity are combined after transformation into Z-scores. A ranking is produced separately for methods in category I and methods in category II. Methods in category II produce substantially higher false positives than methods in category I under phenotype permutation.
We have evaluated the methods ranking stability as a function of several factors that could potentially impact the gene set analysis in different ways, such as the sample size of the microarray datasets, the gene set size, the type of experiment design, and the effect size of the condition under the study. The resulting rankings in the 8 scenarios shown in Table 2 were correlated with the original ranking of the methods (based on all datasets), with the Spearman correlation ranging between 0.78 for large gene sets scenario to 0.98 (all p<0.0001) for unpaired design scenario. The exception to the rule was the paired design scenario for which a 0.34 correlation coefficient was observed with the original ranking. Among the possible factors considered in Table 2, the sample size had the least effect on the methods ranking, with the correlation between the original ranking (based on 42 datasets) and the one based on the smallest and largest 21 datasets being 0.97 and 0.92 respectively (p<0.0001).
Ranking of gene set analysis methods under several scenarios.
| Method | Category | OverallRank | Sample size | Gene set size | Design | Effect | ||||
| Smalln<22 | Largen≥22 | SmallN<66 | LargeN≥66 | Paired | Unpaired | Smallg<24.6% | Largeg≥24.6% | |||
|
| I | 1 | 1 | 4 | 2 | 3 | 12 | 2 | 3 | 3 |
|
| I | 2 | 2 | 1 | 3 | 5 | 6 | 1 | 2 | 4 |
|
| I | 3 | 3 | 2 | 1 | 2 | 1 | 3 | 1 | 1 |
|
| I | 4 | 4 | 3 | 5 | 1 | 2 | 4 | 5 | 2 |
|
| I | 5 | 7 | 5 | 4 | 8 | 8 | 5 | 4 | 6 |
|
| I | 6 | 5 | 8 | 8 | 4 | 5 | 7 | 8 | 8 |
|
| I | 7 | 9 | 6 | 7 | 6 | 11 | 6 | 6 | 11 |
|
| I | 8 | 8 | 7 | 6 | 12 | 4 | 8 | 9 | 5 |
|
| I | 9 | 6 | 10 | 10 | 7 | 9 | 10 | 7 | 9 |
|
| I | 10 | 10 | 9 | 9 | 11 | 10 | 9 | 10 | 7 |
|
| I | 11 | 11 | 11 | 11 | 9 | 3 | 11 | 11 | 10 |
|
| I | 12 | 12 | 12 | 12 | 10 | 7 | 12 | 12 | 12 |
|
| II | 1 | 1 | 1 | 1 | 2 | 1 | 1 | 1 | 1 |
|
| II | 2 | 2 | 2 | 2 | 1 | 2 | 2 | 2 | 2 |
|
| II | 3 | 3 | 3 | 3 | 4 | 3 | 3 | 3 | 3 |
|
| II | 4 | 4 | 4 | 4 | 3 | 4 | 4 | 4 | 4 |
Datasets were divided into small sample size and large sample size using the median of sample sizes, 22, as cut-off. Similarly, the datasets were divided according to the number of genes in the target gene set using the median of gene set sizes, 66, as cut-off. Of the 42 experiments 11 were paired and 31 were not paired designs. To quantify the effect size, datasets were divided into a small effect group and a large effect group based on the % of genes with p<0.05, with the cut-of point being 24.6%.