| Literature DB >> 32560350 |
Michał Burdukiewicz1, Katarzyna Sidorczuk2, Dominik Rafacz1, Filip Pietluch2, Jarosław Chilimoniuk2, Stefan Rödiger3,4, Przemysław Gagat2.
Abstract
Antimicrobial peptides (AMPs) are molecules widespread in all branches of the tree of life that participate in host defense and/or microbial competition. Due to their positive charge, hydrophobicity and amphipathicity, they preferentially disrupt negatively charged bacterial membranes. AMPs are considered an important alternative to traditional antibiotics, especially at the time when multidrug-resistant bacteria being on the rise. Therefore, to reduce the costs of experimental research, robust computational tools for AMP prediction and identification of the best AMP candidates are essential. AmpGram is our novel tool for AMP prediction; it outperforms top-ranking AMP classifiers, including AMPScanner, CAMPR3R and iAMPpred. It is the first AMP prediction tool created for longer AMPs and for high-throughput proteomic screening. AmpGram prediction reliability was confirmed on the example of lactoferrin and thrombin. The former is a well known antimicrobial protein and the latter a cryptic one. Both proteins produce (after protease treatment) functional AMPs that have been experimentally validated at molecular level. The lactoferrin and thrombin AMPs were located in the antimicrobial regions clearly detected by AmpGram. Moreover, AmpGram also provides a list of shot 10 amino acid fragments in the antimicrobial regions, along with their probability predictions; these can be used for further studies and the rational design of new AMPs. AmpGram is available as a web-server, and an easy-to-use R package for proteomic analysis at CRAN repository.Entities:
Keywords: AMP; antimicrobial peptides; host defense peptides; multidrug-resistant bacteria; prediction; proteomic screening; random forest
Mesh:
Substances:
Year: 2020 PMID: 32560350 PMCID: PMC7352166 DOI: 10.3390/ijms21124310
Source DB: PubMed Journal: Int J Mol Sci ISSN: 1422-0067 Impact factor: 5.923
Peptide and protein length distribution in the UniProt [38] and dbAMP [39] database divided into length groups according to the AmpGram benchmark dataset (for details, see Section 3).
| Length Range | UniProt | dbAMP |
|---|---|---|
| 0.85 < 10 | 1119 | 508 |
| 11–19 | 1862 | 1894 |
| 0.8520–26 | 1016 | 1634 |
| 27–36 | 2439 | 1779 |
| 0.8537–60 | 9810 | 2049 |
| 61–710 | 482,852 | 4520 |
| 0.85 > 710 | 45,178 | 5 |
Figure 1Comparison of AmpGram performance with other top-ranking predictors.
Comparison of AmpGram performance with other top-ranking predictors. Programs that do not provide prediction probability are marked with asterisks.
| Software | AUC | Precision | Sensitivity | Specificity |
|---|---|---|---|---|
| AmpGram | 0.9062 | 0.8147 | 0.8543 | 0.8057 |
| ADAM * | 0.7186 | 0.6800 | 0.8259 | 0.6113 |
| 0.85AMPScanner V2 | 0.9641 | 0.9027 | 0.9393 | 0.8988 |
| CAMPR3-ANN * | 0.7854 | 0.7765 | 0.8016 | 0.7692 |
| 0.85CAMPR3-DA | 0.8069 | 0.7286 | 0.8259 | 0.6923 |
| CAMPR3-RF | 0.8958 | 0.7782 | 0.9231 | 0.7368 |
| 0.85CAMPR3-SVM | 0.8363 | 0.7664 | 0.8502 | 0.7409 |
| iAMP-2L * | 0.7895 | 0.8095 | 0.7571 | 0.8219 |
| 0.85iAMPpred (antibacterial) | 0.9008 | 0.8115 | 0.8543 | 0.8016 |
| iAMPpred (antifungal) | 0.9009 | 0.8458 | 0.8219 | 0.8502 |
| 0.85iAMPpred (antiviral) | 0.8397 | 0.7828 | 0.7733 | 0.7854 |
Comparison of AmpGram performance with other top-ranking predictors for 61–710-amino-acid-long AMPs. Programs that do not provide prediction probability are marked with asterisks.
| Software | AUC | Precision | Sensitivity | Specificity |
|---|---|---|---|---|
| AmpGram | 0.8390 | 0.7736 | 0.8542 | 0.7500 |
| ADAM * | 0.6875 | 0.7812 | 0.5208 | 0.8542 |
| AMPScanner V2 | 0.9049 | 0.7963 | 0.8958 | 0.7708 |
| CAMPR3-ANN * | 0.7083 | 0.7000 | 0.7292 | 0.6875 |
| CAMPR3-DA | 0.5221 | 0.5263 | 0.8333 | 0.2500 |
| CAMPR3-RF | 0.6048 | 0.5714 | 0.9167 | 0.3125 |
| CAMPR3-SVM | 0.6228 | 0.5733 | 0.8958 | 0.3333 |
| iAMP-2L * | 0.7292 | 0.9583 | 0.4792 | 0.9792 |
| iAMPpred (antibacterial) | 0.8229 | 0.7188 | 0.9583 | 0.6250 |
| iAMPpred (antifungal) | 0.8110 | 0.7333 | 0.9167 | 0.6667 |
| iAMPpred (antiviral) | 0.7476 | 0.6324 | 0.8958 | 0.4792 |
Figure 2Comparison of AmpGram and AMPscanner [24] performance on the APD and DAMPD dataset with other predictors from Gabere and Noble’s benchmark and according to their methodology [29]. Sequences used to train either AmpGram or AMPScanner were removed from the DAMPD dataset. The benchmark without their removal is presented in Figure S3 in the Supplementary Materials. The very low values of precision are due to the very large negative dataset used (for details, see Section 3).
Comparison of AmpGram and AMPscanner [24] performance on the APD dataset with other predictors from Gabere and Noble’s benchmark and according to their methodology [29]. The very low values of precision are due to the very large negative dataset used (for details, see Section 3).
| Software | AUC | Precision | Sensitivity | Specificity |
|---|---|---|---|---|
| AmpGram | 0.9723 | 0.5531 | 0.9515 | 0.8462 |
| ADAM | 0.8774 | 0.3236 | 0.9095 | 0.6198 |
| AMPA | 0.6394 | 0.4377 | 0.3917 | 0.8994 |
| AMPScanner V2 | 0.9848 | 0.6657 | 0.9743 | 0.9022 |
| CAMPR3-RF | 0.9528 | 0.5337 | 0.9480 | 0.8343 |
| CAMPR3-SVM | 0.9202 | 0.4958 | 0.9060 | 0.8158 |
| DBAASP | 0.7723 | 0.6008 | 0.6281 | 0.9165 |
| MLAMP | 0.8397 | 0.4052 | 0.7560 | 0.7781 |
Comparison of AmpGram and AMPscanner [24] performance on the DAMPD dataset with other predictors from Gabere and Noble’s benchmark and according to their methodology [29]. Sequences used to train either AmpGram or AMPScanner were removed from the dataset. The benchmark without their removal is presented in Table S5 in the Supplementary Materials. The very low values of precision are due to the very large negative dataset used (for details, see Section 3).
| Software | AUC | Precision | Sensitivity | Specificity |
|---|---|---|---|---|
| AmpGram | 0.9321 | 0.3045 | 0.8673 | 0.8472 |
| ADAM | 0.7494 | 0.1540 | 0.7299 | 0.6907 |
| AMPA | 0.6813 | 0.2136 | 0.5355 | 0.8479 |
| AMPScanner V2 | 0.9088 | 0.3661 | 0.8483 | 0.8867 |
| CAMPR3-RF | 0.8162 | 0.1991 | 0.8815 | 0.7265 |
| CAMPR3-SVM | 0.7862 | 0.1926 | 0.8626 | 0.7210 |
| DBAASP | 0.5165 | 0.1014 | 0.1043 | 0.9287 |
| MLAMP | 0.6833 | 0.1695 | 0.4692 | 0.8227 |
Figure 3Distribution of 10-mers along the lactoferrin (A) and thrombin (B) sequences. AMP and non-AMP 10-mers were indicated by black and gray horizontal lines, respectively. The red line represents the cut-off value of 0.5. The red bars mark the fragments that have already been experimentally verified as AMPs: 1–11, 17–41 and 268–284 for lactoferrin [32] and 527–622, 597–622 and 604–622 for thrombin [37]; the sequence coordinates for lactoferrin do not include an N-terminal signal peptide (1–19).
Figure 4Schematic representation of datasets preparation (A), n-grams (B) and decision-making procedure in AmpGram (C). The positive dataset was constructed from sequences downloaded from the dbAMP database [39] (red, green and blue horizontal lines). To create the negative dataset, non-antimicrobial sequences (grey horizontal lines) were retrieved from the UniProt database [38]. The sequences were first concatenated into one string (grey horizontal line), and then cut (black vertical lines) into blocks corresponding in length to sequences from the positive dataset (red, green and blue horizontal line). The extracted blocks were next cut (not indicated in the figure) into subsets corresponding in length to sequences from the positive dataset (red, green and blue circles) and from them individual sequences were randomly selected for the negative dataset (A). Exemplary n-grams used to train AmpGram: the positive n-grams are shaded in red, green and blue, and the negative ones in grey (B). To make a prediction, AmpGram first divides a peptide into subsequences of 10 amino acids (10-mers). For each 10-mer, AmpGram makes a prediction if it is an AMP (true) or not (false) (first model). To scale the prediction for 10-mers to the whole peptide, a lot of statistics is calculated and on their basis AmpGram makes the final prediction (second model). Abbreviations of the statistics: Fraction_true--fraction of positive 10-mers, pred_mean–mean value of prediction, pred_median–median of prediction, n_peptide - number of 10-mers in a peptide, n_pos–number of positive 10-mers, pred_min–minimum value of prediction, pred_max–maximum value of prediction, longest_pos–the longest stretch of consecutively occurring 10-mers predicted as positive, n_pos_10–number of streches comprising of at least 10 10-mers predicted as positive, frac_0_0.2--fraction of 10-mers with prediction in range [0, 0.2], frac_0.2_0.4–fraction of 10-mers with prediction in range (0.2, 0.4], frac_0.4_0.6–fraction of 10-mers with prediction in range (0.4, 0.6], frac_0.6_0.8–fraction of 10-mers with prediction in range (0.6, 0.8], frac_0.8_1–fraction of 10-mers with prediction in range (0.8, 1]) (C).