| Literature DB >> 24931991 |
Yichao Zhou1, Wei Xu1, Bruce R Donald2, Jianyang Zeng1.
Abstract
MOTIVATION: Structure-based computational protein design (SCPR) is an important topic in protein engineering. Under the assumption of a rigid backbone and a finite set of discrete conformations of side-chains, various methods have been proposed to address this problem. A popular method is to combine the dead-end elimination (DEE) and A* tree search algorithms, which provably finds the global minimum energy conformation (GMEC) solution.Entities:
Mesh:
Year: 2014 PMID: 24931991 PMCID: PMC4058937 DOI: 10.1093/bioinformatics/btu264
Source DB: PubMed Journal: Bioinformatics ISSN: 1367-4803 Impact factor: 6.937
Fig. 1.Flow chart of our GA* search algorithm for accelerating protein design. Symbols r represent all parallel expanded rotamers, and p is the total number of expanded nodes. A shaded and rounded square represents a global state, which can be regarded as a global synchronization point. The directional black edges mean that the procedure needs to be done between two synchronization points. The dashed arrows and the double bold arrows represent the data flow among different states and the priority queues, respectively. A group of similar arrows means that the operations are performed in parallel
Fig. 2.Diagram of the GPU states
The comparison results about time efficiency of our parallel against original versions of A* search for protein design
| PDB | Space | OSPREY | A*1 | GA*768 | GA*4992 |
|---|---|---|---|---|---|
| 2·1017 | 21 551 916 | 51 091 | 3075 | 1146 | |
| 2·1014 | 247 585 | 2990 | 296 | 121 | |
| 7·1013 | 96 990 | 1406 | 138 | 73 | |
| 6·1012 | 88 135 | 1771 | 182 | 79 | |
| 3·1014 | 77 614 | 1078 | 99 | 53 | |
| 8·1012 | 64 187 | 1154 | 149 | 57 | |
| 9·1013 | 18 457 | 307 | 33 | 24 | |
| 7·1011 | 8151 | 88 | 18 | 16 | |
| 2·1013 | 6806 | 89 | 18 | 15 | |
| 2·1014 | 6018 | 107 | 18 | 21 |
Notes: Time was measured in millisecond. The results were sorted by the running time needed by OSPREY and only the 10 largest cases are listed here. aThe second column, labeled with ‘Space’, reports the size of conformation search space after the rotamer pruning using iMinDEE. bThe third column, labeled with ‘OSPREY’, reports the running time of the original A* algorithm in OSPREY implemented in Java. cThe fourth column, labeled with ‘A*1’, reports the running time of our new implementation of a single-thread A* algorithm written in C programming language running on a CPU, which adopted the improved computation of heuristic functions, as described in Section 2.2.1. dThe fifth and sixth columns, labeled with ‘GA*768’ and ‘GA*4992’, respectively, report the running time of two fully parallelized A* algorithms running on a GPU, whose numbers of parallel priority queues are 768 and 4992, respectively.
The comparison results about memory consumption of our parallel against original versions of A* search for protein design
| PDB | Space | A*1 | GA*768 | GA*4992 |
|---|---|---|---|---|
| 2·1017 | 31 589 690 | 32 825 074 | 35 517 854 | |
| 2·1014 | 2 910 324 | 3 325 654 | 4 419 100 | |
| 7·1013 | 1 919 055 | 2 282 986 | 3 486 684 | |
| 6·1012 | 1 713 636 | 2 196 315 | 2 960 752 | |
| 3·1014 | 966 196 | 1 255 899 | 1 893 701 | |
| 8·1012 | 1 378 633 | 1 686 558 | 2 354 910 | |
| 9·1013 | 325 634 | 529 810 | 981 302 | |
| 7·1011 | 121 920 | 260 825 | 737 328 | |
| 2·1013 | 129 767 | 211 003 | 618 794 | |
| 2·1014 | 117 053 | 244 399 | 837 359 |
Note: Each column has the same meaning as that in Table 1 except that the numbers in last three columns represent the numbers of expanded nodes in different programs.
Performance of GSMA* with 768 parallel priority queues on 6 test datasets
| PDB | |||||||
|---|---|---|---|---|---|---|---|
| No. of mutable residues | 16 | 18 | 14 | 15 | 15 | 15 | |
| Conformation space | 2·1022 | 2·1020 | 2·1015 | 2·1023 | 3·1020 | 6·1818 | |
| GA*768 search space | 4·107 | 8·106 | 8·106 | 4·107 | 4·107 | 3·107 | |
| 3 × 104 nodes limit | Scan count | 252 | 104 | 99 | 202 | 182 | 109 |
| GMEC gotten | NO | YES | YES | NO | YES | NO | |
| GMEC assured | NO | NO | NO | NO | NO | NO | |
| Correctness | 4% | 100% | 20% | 12% | 32% | 6% | |
| Recovery ratio | 62% | 75% | 85% | 48% | 46% | 48% | |
| 3 × 105 nodes limit | Scan count | 139 | 43 | 36 | 103 | 97 | 55 |
| GMEC gotten | YES | YES | YES | YES | YES | YES | |
| GMEC assured | NO | YES | YES | NO | NO | NO | |
| Correctness | 100% | 100% | 100% | 100% | 100% | 44% | |
| Recovery ratio | 74% | 75% | 87% | 46% | 48% | 54% | |
| 3 × 106 nodes limit | Scan count | 22 | 3 | 3 | 24 | 22 | 18 |
| GMEC gotten | YES | YES | YES | YES | YES | YES | |
| GMEC assured | YES | YES | YES | YES | YES | YES | |
| Correctness | 100% | 100% | 100% | 100% | 100% | 100% | |
| Recovery ratio | 74% | 75% | 87% | 46% | 48% | 53% | |
| 3 × 107 nodes limit | Scan count | 1 | 0 | 0 | 1 | 1 | 1 |
| GMEC gotten | YES | YES | YES | YES | YES | YES | |
| GMEC assured | YES | YES | YES | YES | YES | YES | |
| Correctness | 100% | 100% | 100% | 100% | 100% | 100% | |
| Recovery ratio | 74% | 75% | 87% | 46% | 48% | 53% |
Note: The meaning of each row is explained either in the text or here. aThe row labeled with ‘GA*768 search space’ represents the number of nodes expanded by GA*768 for calculating the best 50 solutions. bThe rows labeled with ‘Scan Count’ represent the number of times that the system ran out of memory, in which a series of operations described in Section 2.5 were executed.