| Literature DB >> 25670805 |
Yoshinori Fukasawa1, Junko Tsuji1, Szu-Chin Fu1, Kentaro Tomii2, Paul Horton3, Kenichiro Imai4.
Abstract
Mitochondria provide numerous essential functions for cells and their dysfunction leads to a variety of diseases. Thus, obtaining a complete mitochondrial proteome should be a crucial step toward understanding the roles of mitochondria. Many mitochondrial proteins have been identified experimentally but a complete list is not yet available. To fill this gap, methods to computationally predict mitochondrial proteins from amino acid sequence have been developed and are widely used, but unfortunately, their accuracy is far from perfect. Here we describeEntities:
Mesh:
Substances:
Year: 2015 PMID: 25670805 PMCID: PMC4390256 DOI: 10.1074/mcp.M114.043083
Source DB: PubMed Journal: Mol Cell Proteomics ISSN: 1535-9476 Impact factor: 5.911
Fig. 1.Local sequences and prediction performance of cleavage sites. A, Sequence logo of MPP cleavage sites partitioned into three classes (MPP only, MPP+Icp55, MPP+Oct1) based on recent proteomics data. The dashed line boxes show the range of positions covered by the PWMs for MPP, Oct1, and Icp55. B, Cleavage site accuracy comparison on the yeast data set. Error bars show the standard error of mean estimation based on 10-fold cross validation (only MitoFates is retrained, the other tools are used as distributed without retraining but their prediction accuracy still varies between test folds).
Fig. 2.MTS discrimination performance comparison between MitoFates and previous predictors on an independent test data set. A, Comparison by PR-curve. B, True positive rate versus false positive rate. C, Statistical significance (vertical axis) of the true positive rate difference between MitoFates and other predictors plotted against false positive rate. For each input sequence the predictors output both a score (for TPpred2 we extracted the GRHCRF-scores from their software) and a label (mitochondrial, ER, or other); the dashed lines show performance based purely on the scores, and the solid lines always count mislabeled mitochondrial proteins as false negatives and nonmitochondrial proteins with nonmitochondrial labels as true negatives, regardless of score.
Comparison of ROC AUC and MCC on an independent test data set
| Predictor | ROC AUC | MCC |
|---|---|---|
| MitoFates (cutoff: 0.5) | 0.954 | 0.465 |
| MitoFates (cutoff: 0.385) | 0.446 | |
| TPpred2 | 0.948 | 0.355 |
| Predotar | 0.939 | 0.304 |
| TargetP | 0.933 | 0.242 |
| MitoProtII | 0.941 | 0.217 |
Fig. 3.Positively charged amphiphilicity (PA) score and presequence specific motifs. A, Histograms of the distribution of maximum hydrophobic moment score (above) and PA score (below) in the N-terminal 30 residues of proteins with and without a presequence. B, The top six presequence specific hexamers are listed with their statistical significance, coverage in the positive, and negative training examples, PA score, and a sequence logo depicting their matches in the positive data. Arrows show the hydrophobic moments. Blue and gray indicate higher or lower PA score than the 90th percentile score over all hexamers in the positive and negative examples, respectively.
Fig. 4.The three clusters of yeast presequences. A, Presequence feature vectors, colored by cluster, shown as mapped to three dimensions by PCA. B, Box plots of the nine features used for clustering are shown for each cluster.
Fig. 5.An example of prediction output from the MitoFates web server is shown.