| Literature DB >> 17884914 |
Masatoshi Ichihara1, Yoshiki Murakumo, Akio Masuda, Toru Matsuura, Naoya Asai, Mayumi Jijiwa, Maki Ishida, Jun Shinmi, Hiroshi Yatsuya, Shanlou Qiao, Masahide Takahashi, Kinji Ohno.
Abstract
We developed a simple algorithm, i-Score (inhibitory-Score), to predict active siRNAs by applying a linear regression model to 2431 siRNAs. Our algorithm is exclusively comprised of nucleotide (nt) preferences at each position, and no other parameters are taken into account. Using a validation dataset comprised of 419 siRNAs, we found that the prediction accuracy of i-Score is as good as those of s-Biopredsi, ThermoComposition21 and DSIR, which employ a neural network model or more parameters in a linear regression model. Reynolds and Katoh also predict active siRNAs efficiently, but the numbers of siRNAs predicted to be active are less than one-eighth of that of i-Score. We additionally found that exclusion of thermostable siRNAs, whose whole stacking energy (DeltaG) is less than -34.6 kcal/mol, improves the prediction accuracy in i-Score, s-Biopredsi, ThermoComposition21 and DSIR. We also developed a universal target vector, pSELL, with which we can assay an siRNA activity of any sequence in either the sense or antisense direction. We assayed 86 siRNAs in HEK293 cells using pSELL, and validated applicability of i-Score and the whole DeltaG value in designing siRNAs.Entities:
Mesh:
Substances:
Year: 2007 PMID: 17884914 PMCID: PMC2094068 DOI: 10.1093/nar/gkm699
Source DB: PubMed Journal: Nucleic Acids Res ISSN: 0305-1048 Impact factor: 16.971
Figure 5.In vitro validation of i-Score using pSELL and pDual vectors. (A) The pSELL target vector carries the SV40 promoter, the EGFP gene, the target-cloning site, IRES and the firefly luciferase gene. Synthesized oligonucleotides should carry HindIII- and BamHI/BglII-competent restriction sites at both ends. pSELL can accommodate the synthesized oligonucleotides in either direction. The pDual effector vector harbors the U1 and H1 promoters in opposite directions, and the cloning site between them (30). pDual accommodates the same oligonucleotides as pSELL in the sense direction. siRNA synthesized by pDual works on the target mRNA synthesized by pSELL. We can thus quantify the siRNA activity by measuring the activity of either EGFP or the firefly luciferase. Scattered plots of 86 observed and predicted siRNA activities above (B) and below (C) the whole ΔG value of −34.6 kcal/mol. The correlation coefficient of thermodynamically unstable siRNAs (B) is better than that of thermostable siRNAs (C).
Figure 1.Definition of i-Score (A) and scoring parameters at each position (B). Parameters are normalized to give the best and the worst i-Scores of 100 and 0, respectively. The average score at each position is 2.60 (dotted line). Scores above 2.60 indicate preferred nucleotides. ‘P’ is a probability score of nucleotide ‘n’ at position ‘m’ on the antisense strand (Supplementary Table 1).
Figure 2.Observed siRNA activities in dataset B are plotted against predicted siRNA activities by i-Score (A), s-Biopredsi (B), ThermoComposition21 (C) and DSIR (D). ‘R’ values represent the Pearson correlation coefficients, which are also indicated in Table 1. (E) ROC curves of the four algorithms. Areas under the curves (AUC) of i-Score, s-Biopredsi, ThermoComposition21 and DSIR are 0.776 (95% confidence interval, 0.732–0.820), 0.770 (0.726–0.814), 0.795 (0.753–0.837) and 0.781 (0.738–0.825), respectively. There are no statistical differences of AUCs among the four algorithms.
Pearson correlation coefficients between observed and predicted siRNA activities by four second-generation algorithms
| Dataset A | Dataset B | |
|---|---|---|
| 0.635 | 0.557 | |
| 0.665 | 0.546 | |
| 0.635 | 0.577 | |
| 0.687 | 0.554 |
Dataset A is used to train all the algorithms, whereas dataset B is not used as a training dataset in any algorithms.
Comparison of eight algorithms using subset A831
| Threshold | % of effective siRNAs | No (%) of siRNAs matching the threshold | Pearson correlation coefficient |
|---|---|---|---|
| 0.592 | |||
| ≥65.9 | 90 | 72 (8.8%) | |
| ≥63.0 | 80 | 117 (14.1%) | |
| ≥59.4 | 75 | 166 (20.0%) | |
| 0.618 | |||
| ≥0.807 | 90 | 73 (8.8%) | |
| ≥0.767 | 80 | 141 (17.0%) | |
| ≥0.734 | 75 | 197 (23.7%) | |
| N.D. | |||
| ≥9 | 88.9 | 9 (1.1%) | |
| ≥8 | 80.4 | 46 (5.5%) | |
| ≥7 | 71.0 | 107 (12.9%) | |
| N.D. | |||
| Ia | 73.4 | 64 (7.7%) | |
| Ia + Ib | 68.0 | 125 (15.0%) | |
| N.D. | |||
| ≥5 | 81.0 | 21 (2.5%) | |
| ≥4 | 61.3 | 80 (9.6%) | |
| ≥3 | 64.8 | 179 (21.5%) | |
| 0.427 | |||
| ≥ 101.1 | 90 | 6 (0.7%) | |
| ≥ 87.5 | 80 | 31 (3.7%) | |
| ≥79.5 | 75 | 80 (11.1%) | |
| N.D. | |||
| =4 | 50.0 | 2 (0.2%) | |
| ≥3 | 51.2 | 41 (4.9%) | |
| ≥2 | 52.7 | 184 (22.1%) | |
| 0.174 | |||
| n.a. | 90 | n.a. | |
| n.a. | 80 | n.a. | |
| ≥17.1 | 75 | 4 (0.5%) |
Note that i-Score1600, s-Boipredsi1600, Reynolds and Katoh predict active siRNAs with ∼90% accuracy, but the chances of predicting such active siRNAs in a given mRNA with Reynolds and Katoh are ∼1/8 and ∼1/12, respectively, of those with i-Score1600 and s-BoiPredSi1600. ThermoComposition21 and DSIR are not included in this analysis, because we cannot avoid overfitting for these two algorithms for subset A831.
aRatios of experimentally proved active siRNAs among siRNAs predicted to be active according to variable thresholds of eight different algorithms. The experimentally proved active siRNAs are arbitrarily defined to those suppressing the gene expression levels to less than 25% of a control.
bThe i-Score1600 and s-Biopredisi1600 scores are different from i-Score and s-Biopredsi. In order to avoid overfitting for these two algorithms, the modeling algorithms of i-Score and s-Biopredsi are applied to subset A1600 to calculate scoring parameters for i-Score1600 and s-Biopredisi1600, respectively. The Pearson correlation coefficients between i-Score and i-Score1600, and s-Biopredsi and s-Biopredsi1600 are 0.990 and 0.985, respectively.
cAs i-Score1600, s-Biopredsi1600, Katoh and Takasaki are continuous numeric scores, thresholds are arbitrarily set so that 90, 80 and 75% siRNAs above the threshold are experimentally proved active. For example, for i-Score1600, all the 831 siRNAs in subset A831 are sorted in descending order of i-Score1600. The ratio of experimentally proved active siRNAs decreases with decreasing i-Score1600. When the lower limit of i-Score1600 is set to 65.9, 72 siRNAs are included in this category and 65 suppress the gene expression levels to less than 25% of a control. This is how we set an i-Score1600 threshold for 90% (65/72). The 72 siRNAs comprise 8.8% of the 831 siRNAs in subset A831. This indicates that when we synthesize an siRNA with i-Score1600 of 65.9 or higher, we can expect that the chance of obtaining an active siRNA is 90%.
dFor algorithms yielding ordinal or nominal numbers, the indicated ranks are used as the thresholds. For example, for Reynolds, the number of siRNAs equal to or higher than the score of 9 is 9, which comprises 1.1% of the 831 siRNAs, and 8 (88.9%) of the 9 siRNAs suppress gene expression levels to less than 25% of a control.
eNo subgroup of siRNAs matches the criteria of ≥90% or ≥80% prediction accuracy.
Figure 3.Contours of the whole ΔG values overlaid on the scatted plots of observed and predicted siRNA activities with (A) or without (B) dots indicating individual siRNAs for dataset B. Area H indicates siRNAs for which i-Score successfully predicts siRNA activities, whereas area L indicates siRNAs for which i-Score falsely predicts inactive siRNAs. Thermodynamically unstable siRNAs with high whole ΔG values tend to align on a linear regression line (data not shown) going from the bottom left up to the top right, and form a cluster of well-predicted siRNAs (area H) at the top right corner. On the other hand, thermostable siRNAs stay away from the regression line, and form a cluster of poorly predicted siRNAs (area L) at the top left corner. The plots are drawn by the ‘contour plot’ functionality of the JMP-IN software. (C) Correlation coefficients between the observed and predicted siRNA activities in subgroups of siRNAs whose whole ΔG values are equal to or more than the indicated values. For example, there are 101 siRNAs whose whole ΔG values are equal to or more than −34.6 kcal/mol. The correlation coefficient of the 101 siRNAs with i-Score is 0.723, and hence 0.723 is plotted on –34.6 (arrow). This is also illustrated in Figure 4. The correlation coefficient tends to go down when the whole ΔG threshold further goes higher, likely because we can include a limited number of siRNAs in the analysis.
Figure 4.Scattered plots of observed and predicted siRNAs categorized by the whole ΔG values for datasets B. The threshold of whole ΔG values is indicated on top of each panel. i-Score predicts activities of thermodynamically unstable siRNAs (A) more accurately than those of thermostable siRNAs (B). (C) ROC curves of thermodynamically unstable (≥−34.6 kcal/mol) and stable (<−34.6 kcal/mol) siRNAs in dataset B using i-Score. AUCs of unstable and stable siRNAs are 0.882 (95% confidence interval, 0.814–0.950) and 0.750 (0.697–0.803), respectively. AUC of the whole dataset B is 0.776 (0.732–0.820).
Genome-wide prediction of active siRNAs by i-Score
| Whole Δ | siRNAs/kb | siRNAs/mRNA | Best | 10th | |
|---|---|---|---|---|---|
| >65 | – | 82.9 | 132 (99.8%) | 82.6 ± 4.3 | 75.6 ± 4.7 |
| >65 | >−34 | 50.5 | 65 (95.9%) | 82.2 ± 5.3 | 74.2 ± 7.2 |
| >65 | >−30 | 26.7 | 33 (89.2%) | 81.8 ± 5.9 | 72.8 ± 7.8 |
| >70 | – | 34.0 | 53 (98.8%) | – | – |
| >70 | >−34 | 24.2 | 33 (94.0%) | – | – |
| >70 | >−30 | 14.1 | 19 (86.9%) | – | – |
Average number of active siRNAs per kb of the human RefSeq mRNAs. The NCBI RefSeq Database Build 35.1 includes 40 768 mRNAs, and we analyzed each alternatively spliced transcript as an independent mRNA. The total number of nucleotides that we analyzed is 100 738 984, and the mean and SD of the mRNA length is 2471 ± 2039 bases.
Median number of active siRNAs per mRNA in the human RefSeq database. As the numbers of active siRNAs do not follow a Gaussian distribution, median numbers are represented. The number in parenthesis indicates a percentage of mRNAs among the 40 768 transcripts, harboring at least one siRNA meeting the indicated criteria.
Mean and SD of the highest i-Score for each mRNA. mRNAs with no siRNAs above the indicated whole ΔG threshold are excluded from the analysis.
Mean and SD of the 10th highest i-Score for each mRNA. mRNAs with less than 10 siRNAs above the indicated whole ΔG threshold are excluded from the analysis.
No threshold is placed for the whole ΔG.
Because means of the best and the 10th highest i-Scores are independent of the i-Score threshold, values are indicated only on rows of i-Score >65.