| Literature DB >> 27688811 |
Wajdi Dhifli1, Abdoulaye Baniré Diallo1.
Abstract
BACKGROUND: Studying the functions and structures of proteins is important for understanding the molecular mechanisms of life. The number of publicly available protein structures has increasingly become extremely large. Still, the classification of a protein structure remains a difficult, costly, and time consuming task. The difficulties are often due to the essential role of spatial and topological structures in the classification of protein structures.Entities:
Keywords: Graph classification; Protein 3D-structure; Protein classification
Year: 2016 PMID: 27688811 PMCID: PMC5034655 DOI: 10.1186/s13040-016-0108-2
Source DB: PubMed Journal: BioData Min ISSN: 1756-0381 Impact factor: 2.522
Fig. 1The human hemoglobin protein 3D-structure (PDBID: 1GZX) and its corresponding graph representation. Nodes and edges represent, respectively, amino acids from the structure and links between them. Blue edges represent links from the primary structure and gray edges are spatial links between distant amino acids
Characteristics of the experimental datasets
| Dataset | SCOP ID | Family name | Pos. | Neg. | Avg. ∣ | Avg. ∣ | Max. ∣ | Max. ∣ |
|---|---|---|---|---|---|---|---|---|
| DS1 | 48623 | Vertebrate phospholipase A2 | 29 | 29 | 160 | 628 | 451 | 1812 |
| DS2 | 52592 | G-proteins | 33 | 33 | 246 | 971 | 897 | 3544 |
| DS3 | 48942 | C1-set domains | 38 | 38 | 238 | 928 | 768 | 2962 |
| DS4 | 56437 | C-type lectin domains | 38 | 38 | 185 | 719 | 775 | 3016 |
| DS5 | 56251 | Proteasome subunits | 35 | 35 | 231 | 929 | 897 | 3544 |
| DS6 | 88854 | Protein kinases, catalyc subunits | 41 | 41 | 275 | 1077 | 775 | 3016 |
SCOP ID, Family name, Pos., Neg., Avg. ∣V∣, Avg. ∣E∣, Max. ∣V∣ and Max. ∣E∣ correspond respectively to the identifier of the positive protein family in SCOP, its name, the number of positive examples, the number of negative examples, the average number of nodes, the average number of edges, the maximal number of nodes and the maximal number of edges in each dataset
Fig. 2Classification accuracy of PROTNN using different distance measures
Fig. 3Tendancy of the average accuracy of PROTNN over the six datasets for k∈ [1,10]. The dashed line represents the linear tendancy of the results
Empirical ranking of the structural and topological attributes
| Data | Attributes | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| A1 | A2 | A3 | A4 | A5 | A6 | A7 | A8 | A9 | A10 | A11 | A12 | A13 | A14 | A15 | A16 | A17 | A18 | |
| DS1 | 0 | 2 | 1 | 6 | 2 | 1 | 5 | 9 | 10 | 1 | 2 | 8 | 6 | 13 | 16 | 17 | 12 | 17 |
| DS2 | 8 | 12 | 15 | 16 | 18 | 4 | 9 | 16 | 23 | 17 | 9 | 21 | 11 | 14 | 25 | 17 | 23 | 9 |
| DS3 | 8 | 13 | 2 | 6 | 17 | 10 | 16 | 11 | 11 | 4 | 8 | 18 | 21 | 2 | 21 | 23 | 9 | 18 |
| DS4 | 4 | 7 | 21 | 17 | 20 | 6 | 11 | 17 | 16 | 7 | 2 | 14 | 21 | 22 | 20 | 21 | 24 | 17 |
| DS5 | 12 | 12 | 8 | 10 | 12 | 5 | 7 | 7 | 17 | 17 | 7 | 23 | 23 | 9 | 20 | 9 | 19 | 18 |
| DS6 | 5 | 11 | 9 | 8 | 11 | 6 | 14 | 14 | 13 | 6 | 1 | 17 | 14 | 18 | 24 | 10 | 17 | 13 |
| Total | 37 | 57 | 56 | 63 | 80 | 32 | 62 | 74 | 90 | 52 | 29 |
|
| 78 | 126 |
|
| 92 |
| Score | 0.25 | 0.38 | 0.37 | 0.42 | 0.53 | 0.21 | 0.41 | 0.49 | 0.6 | 0.35 | 0.19 |
|
| 0.52 |
|
|
| 0.61 |
| Rank | 16 | 13 | 14 | 11 | 8 | 17 | 12 | 10 | 7 | 15 | 18 |
|
| 9 |
|
|
| 6 |
The boldface numbers highlight the best performance
Accuracy comparison of PROTNN and PROTSVM
| Dataset | Classification approach | ||
|---|---|---|---|
|
| ProtSVM(linear) | ProtSVM(rbf) | |
| DS1 |
| 0.88 | 0.83 |
| DS2 |
| 0.68 | 0.56 |
| DS3 |
| 0.87 | 0.78 |
| DS4 |
| 0.80 | 0.82 |
| DS5 |
| 0.79 | 0.73 |
| DS6 |
| 0.84 | 0.72 |
| Avg. accuracy1 |
| 0.81 ±0.07 | 0.74 ±0.1 |
1Average classification accuracy of each classification approach over the six datasets. The boldface numbers highlight the best performance
Accuracy comparison of PROTNN with other classification techniques
| Dataset | Classification approach | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Blast | Sheba | FatCat | CE | LPGBCMP | D&D | GAIA |
|
| |
| DS1 | 0.88 | 0.81 |
| 0.45 | 0.88 | 0.93 |
| 0.97 | 0.97 |
| DS2 | 0.82 | 0.86 |
| 0.49 | 0.73 | 0.76 | 0.66 | 0.8 |
|
| DS3 | 0.9 | 0.95 | 0.84 | 0.59 | 0.90 | 0.96 | 0.89 | 0.96 |
|
| DS4 | 0.76 | 0.92 |
| 0.46 | 0.9 | 0.93 | 0.89 | 0.97 | 0.97 |
| DS5 | 0.86 |
| 0.94 | 0.76 | 0.87 | 0.89 | 0.72 | 0.9 | 0.94 |
| DS6 | 0.78 |
| 0.94 | 0.81 | 0.91 | 0.95 | 0.87 | 0.96 | 0.96 |
| Avg. accuracy1 | 0.83 ±0.05 | 0.92 ±0.07 | 0.94 ±0.06 | 0.59 ±0.15 | 0.86 ±0.06 | 0.9 ±0.07 | 0.84 ±0.12 | 0.93 ±0.06 |
|
| Avg. distances2 | 0.14 ±0.07 | 0.05 ±0.07 | 0.04 ±0.05 | 0.38 ±0.15 | 0.11 ±0.03 | 0.7 ±0.04 | 0.14 ±0.09 | 0.05 ±0.03 |
|
| Rank | 8 | 4 | 2 | 9 | 6 | 5 | 7 | 3 |
|
1Average classification accuracy of each classification approach over the six datasets
2Average of the distances between the accuracy of each approach and the best obtained accuracy with each dataset
The boldface numbers highlight the best performance
Fig. 4Runtime comparison in log-scale of PROTNN and FatCat. The runtime of PROTNN is separated for the main steps. ProtNN(all) is the sum of its three steps: Graph transformation, Attributes computation and Classification
Runtime results of PROTNN, FatCat and CE on the entire Protein Data Bank
| Task | Total runtime1 | Runtime1/protein |
|---|---|---|
| Building graph models | 23h:9m:57s | 0.9s |
| Computation of attributes | 5d:8h:12m:29s | 4.9s |
| Classification | 2h:55m:15s | 0.1s |
|
| 6d:10h:17m:41s | 5.9s |
|
| Forever2 | 1d:18h:31m:35s3 |
|
| Forever2 | 1d:8h:37m:34s3 |
1The runtime is expressed in terms of days:hours:minutes:seconds
2The program did not finish running within two weeks
3The average runtime of randomly selected 100 proteins