| Literature DB >> 31207991 |
Michele Montaruli1, Domenico Alberga2, Fulvio Ciriaco3, Daniela Trisciuzzi4, Anna Rita Tondo5, Giuseppe Felice Mangiatordi6, Orazio Nicolotti7.
Abstract
In this continuing work, we have updated our recently proposed Multi-fingerprint Similarity Search algorithm (MuSSel) by enabling the generation of dominant ionized species at a physiological pH and the exploration of a larger data domain, which included more than half a million high-quality small molecules extracted from the latest release of ChEMBL (version 24.1, at the time of writing). Provided with a high biological assay confidence score, these selected compounds explored up to 2822 protein drug targets. To improve the data accuracy, samples marked as prodrugs or with equivocal biological annotations were not considered. Notably, MuSSel performances were overall improved by using an object-relational database management system based on PostgreSQL. In order to challenge the real effectiveness of MuSSel in predicting relevant therapeutic drug targets, we analyzed a pool of 36 external bioactive compounds published in the Journal of Medicinal Chemistry from October to December 2018. This study demonstrates that the use of highly curated chemical and biological experimental data on one side, and a powerful multi-fingerprint search algorithm on the other, can be of the utmost importance in addressing the fate of newly conceived small molecules, by strongly reducing the attrition of early phases of drug discovery programs.Entities:
Keywords: data quality; molecular similarity; multi-fingerprint; protein drug target prediction
Mesh:
Substances:
Year: 2019 PMID: 31207991 PMCID: PMC6631269 DOI: 10.3390/molecules24122233
Source DB: PubMed Journal: Molecules ISSN: 1420-3049 Impact factor: 4.411
Fingerprint notations along with the open-source software packages used for their calculation.
| Fingerprints Name | Description | Package | Reference |
|---|---|---|---|
|
| Morgan connectivity invariants ( | RDKit | [ |
|
| Morgan feature invariants ( | RDKit | [ |
|
| Atom pairs fingerprint | RDKit | [ |
|
| SMARTS Pattern fingerprint | RDKit | [ |
|
| Daylight-like topological fingerprint | RDKit | [ |
|
| Topological torsion fingerprint | RDKit | [ |
|
| Indexes linear fragments up to 7 atoms | Pybel | [ |
|
| Pubchem fingerprints | CDK | [ |
|
| CDK | [ | |
|
| CDK | [ | |
|
| Graph fingerprint which does not take bond orders into account | CDK | [ |
|
| Bit set type fingerprint based on 307 substructures | CDK | [ |
|
| Fingerprint based on hybridization state of atoms | CDK | [ |
Figure 1Similarity comparisons of one million neutral vs. ionized pairs of compounds by using klekota_roth, cdk_maccs, pubchem and substructure, and FeatMFP1 FPs. Orange/purple pairs have similarity values always under/over the threshold, respectively, irrespective of the ionization state. Green/red pairs have similarity values awarded/penalized after ionization, respectively.
For both the K and IC50 pools, predictions are based on first using the neutral database and then the ionized database on the same external dataset comprised of 5000 ionized compounds at a physiological pH randomly discarded by the training set based on ChEMBL (version 24.1). Using both K and IC50 protein drug target data, the predictions were considered successful if a match was found as the top-one (p1) or within the top-five (p5).
|
|
| |||
|---|---|---|---|---|
|
|
|
|
| |
| Neutral database | 89.72% | 92.82% | 86.80% | 90.20% |
| Ionized database | 91.08% | 93.16% | 88.72% | 92.24% |
1 The calibration parameters were kept unchanged, as in our previous study [2].
Each of Ext1, Ext2, and Ext3 comprised 300 molecules randomly taken from the difference between ChEMBL (version 23) and ChEMBL (version 22.1). Ext4 comprised 1000 compounds randomly discarded from the training set based on ChEMBL (version 24.1). Using both K and IC50 protein drug target data, the predictions were considered successful if a match was found as the top-one (p1) or within the top-five (p5) targets.
|
|
| |||
|---|---|---|---|---|
|
|
|
|
| |
| Ext1 ( | 90.67% | 96.00% (56.20%) * | 88.00% | 93.33% (35.00%) * |
| Ext2 ( | 90.33% | 96.00% (48.60%) * | 92.00% | 95.00% (31.70%) * |
| Ext3 ( | 93.67% | 97.33% (51.40%) * | 89.33% | 92.00% (29.30%) * |
| Ext4 ( | 90.77% | 94.32% | 90.10% | 93.20% |
1 The calibration parameters were kept unchanged, as in our previous study [2]. * For the ease comparison, the p5 values obtained in our previous study [2] are reported in parentheses.
Chemical structures of the 18 entries selected from the Journal of Medicinal Chemistry (i.e., inspecting papers published from October to December 2018) whose protein drug targets were successfully predicted. For each entry, the name of the protein drug target with the corresponding number of associated compounds, as well as the ChEMBL ID available in MuSSel, are reported. A parallel table with the unsuccessful cases is enclosed in the Supplementary Materials.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|