| Literature DB >> 21143806 |
Paul D Yoo1, Yung Shwen Ho, Jason Ng, Michael Charleston, Nitin K Saksena, Pengyi Yang, Albert Y Zomaya.
Abstract
Changes to the glycosylation profile on HIV <span class="Gene">gp120 can influence viral pathogenesis and alter AIDS disease progression. The characterization of glycosylation differences at the sequence level is inadequate as the placement of carbohydrates is structurally complex. However, no structural framework is available to date for the study of HIV disease progression. In this study, we propose a novel machine-learning based framework for the prediction of AIDS disease progression in three stages (RP, SP, and LTNP) using the HIV structural gp120 profile. This new intelligent framework proves to be accurate and provides an important benchmark for predicting AIDS disease progression computationally. The model is trained using a novel HIV gp120 glycosylation structural profile to detect possible stages of AIDS disease progression for the target sequences of HIV+ individuals. The performance of the proposed model was compared to seven existing different machine-learning models on newly proposed gp120-Benchmark_1 dataset in terms of error-rate (MSE), accuracy (CCI), stability (STD), and complexity (TBM). The novel framework showed better predictive performance with 67.82% CCI, 30.21 MSE, 0.8 STD, and 2.62 TBM on the three stages of AIDS disease progression of 50 HIV+ individuals. This framework is an invaluable bioinformatics tool that will be useful to the clinical assessment of viral pathogenesis.Entities:
Mesh:
Substances:
Year: 2010 PMID: 21143806 PMCID: PMC3005921 DOI: 10.1186/1471-2164-11-S4-S22
Source DB: PubMed Journal: BMC Genomics ISSN: 1471-2164 Impact factor: 3.969
Patient dataset used for structural glycan profiling
| Cohort/Country | Sample | GenBank Accession No | Year of sample collection | Disease type |
|---|---|---|---|---|
| USA | A1 | AY835754 | 1982 | RP |
| USA | A2 | AY835765 | 1984 | RP |
| USA | A3 | AY835775 | 1986 | RP |
| USA | B4 | AY835777 | 1983 | RP |
| USA | B6 | AY835778 | 1986 | RP |
| USA | C7 | AY835779 | 1984 | RP |
| USA | C8 | AY835780 | 1986 | RP |
| USA | D9 | AY835781 | 1983 | RP |
| USA | D10 | AY835755 | 1985 | RP |
| USA | D11 | AY835756 | 1986 | RP |
| USA | E12 | AY835757 | 1986 | RP |
| USA | F1 | AY835759 | 1982 | SP |
| USA | F2 | AY835760 | 1987 | SP |
| USA | F3 | AY835761 | 1991 | SP |
| USA | G4 | AY835762 | 1984 | SP |
| USA | G5 | AY835763 | 1988 | SP |
| USA | G6 | AY835764 | 1992 | SP |
| USA | H7 | AY835766 | 1989 | SP |
| USA | H8 | AY835767 | 1993 | SP |
| USA | J1 | AY835769 | 1985 | SP |
| USA | K4 | AY835771 | 1986 | SP |
| USA | K5 | AY835772 | 1992 | SP |
| USA | K6 | AY835773 | 1994 | SP |
| USA | L7 | AY835774 | 1986 | SP |
| USA | M1 | AY835748 | 1983 | LTNP |
| USA | M2 | AY835749 | 1986 | LTNP |
| USA | M3 | AY835750 | 1989 | LTNP |
| USA | M4 | AY835751 | 1990 | LTNP |
| USA | M5 | AY835752 | 1992 | LTNP |
| USA | N8 | AY835753 | 1996 | LTNP |
| Canada | CAN_A_1 | AY779564 | 1996 | LTNP |
| Canada | CAN_A_3 | AY779550 | 1998 | LTNP |
| Canada | CAN_A_5 | AY779551 | 1999 | LTNP |
| Canada | CAN_A_6 | AY779552 | 2000 | LTNP |
| Canada | CAN_B_3 | AY779553 | 1994 | SP |
| Canada | CAN_B_4 | AY779554 | 1997 | SP |
| Canada | CAN_B_5 | AY779555 | 1998 | SP |
| Canada | CAN_B_6 | AY779556 | 1999 | SP |
| Canada | CAN_C_2 | AY779557 | 1992 | SP |
| Canada | CAN_C_3 | AY779558 | 1993 | SP |
| Canada | CAN_C_5 | AY779559 | 1994 | SP |
| Canada | CAN_C_6 | AY779560 | 1994 | SP |
| Canada | CAN_C_8 | AY779561 | 1996 | SP |
| Canada | CAN_C_10 | AY779562 | 1998 | SP |
| Australia | 1181 | GQ995529 | 1995 | SP |
| Australia | 1182 | GQ995528 | 1995 | SP |
| Australia | BB_76 | GQ995532 | 1982 | SP |
| Australia | BB_24 | GQ995530 | 1984 | SP |
| Australia | BB_42 | GQ995531 | 1984 | SP |
| Australia | BB_92 | GQ995533 | 1983 | SP |
Figure 1Annotation of the glycans on the template env gp120 crystal structure PB4C (left-hand), and variations in structural locations of glycans (right-hand).
Figure 2Modular hierarchical kernel experts as a global model. HME has a tree-like structure that uses the divide and conquer principle to learn the interactions between HIV glycosylation sites. At the bottom of the tree (leaf nodes) are multiple locality effective support vector, that will analyse the similarities between the given glycosylation sites.
Figure 3The basic architecture of semi-parameterized SV-based local models. The centroid vectors from voronoi region for each training sample x used in the SVM decision function. SVM is considered as a purely non-parametric model, whereas SV-HMM is considered as semi-parametric model as it adopts the method of grouping the associated input vectors in each class i. RBF kernel has been used for the SV-based local models.
Figure 4The flowchart of SV-HMM showing the stepwise procedure. The above figure shows the stepwise procedure we have performed. (1) data collection, building gp120 benchmark dataset and pre-processing datasets; (2) structural gp120 profile construction including matching up and calculating the 3D distance between every glycan of the query and template models. (3) the information obtained in (2) and (3) were combined and normalised to fall in the interval [--1, 1] to be fed into networks; (4) target levels were assigned to each profile (positive, +1, for RPs, 0 for SP. –1 for LTNP); (5) a hold-out method, to divide the combined dataset into ten subsets (training and testing sets); (6) model training on each set, to create a model; (7) simulation of each model on the test set, to obtain predicted outputs; and (8) post-processing to find predicted HIV progressor groups. The procedure from (6) to (8) was performed iteratively until we obtained the most suitable kernel and the optimal hyperparameters for SV-HMM for gp120 benchmark dataset.
Model comparison on gp120_Benchmark_1 dataset
| Models | MSERP | MSESP | MSELTNP | MSEOverall | CCI | STD | TBM |
|---|---|---|---|---|---|---|---|
| SV-HMM | 38.02 | 30.46 | 45.57 | 30.21 | 67.82 | 0.8 | 2.62 |
| HMESVM | 36.35 | 24.13 | 44.43 | 29.76 | 69.38 | 1.9 | 1.35 |
| SVMLIB | 54.53 | 20.68 | 44.43 | 32.65 | 67.39 | 1.1 | 0.17 |
| DecorateSVM | 81.82 | 10.34 | 55.55 | 32.88 | 65.31 | 2.7 | 3.02 |
| MLP | 72.43 | 27.59 | 55.57 | 27.27 | 57.14 | 4.2 | 4.14 |
| RBFN | 63.63 | 27.58 | 77.78 | 29.32 | 55.10 | 3.8 | 0.06 |
| Logistic | 63.63 | 34.48 | 66.66 | 31.35 | 53.06 | 1.0 | 0.11 |
| J48 | 90.91 | 37.93 | 66.67 | 37.77 | 44.90 | 3.0 | 0.01 |
The final CCI value was calculated based on the average of all the prediction accuracies ob-served during the tenfold validation process. MSE, TBM and STD (ST Deviation obtained from ten sub-samples) indicate the accuracy, complexity and stability of the model respectively.