| Literature DB >> 23777206 |
Xavier I Ambroggio1, Jennifer Dommer, Vivek Gopalan, Eleca J Dunham, Jeffery K Taubenberger, Darrell E Hurt.
Abstract
BACKGROUND: Influenza A viruses possess RNA genomes that mutate frequently in response to immune pressures. The mutations in the hemagglutinin genes are particularly significant, as the hemagglutinin proteins mediate attachment and fusion to host cells, thereby influencing viral pathogenicity and species specificity. Large-scale influenza A genome sequencing efforts have been ongoing to understand past epidemics and pandemics and anticipate future outbreaks. Sequencing efforts thus far have generated nearly 9,000 distinct hemagglutinin amino acid sequences. DESCRIPTION: Comparative models for all publicly available influenza A hemagglutinin protein sequences (8,769 to date) were generated using the Rosetta modeling suite. The C-alpha root mean square deviations between a randomly chosen test set of models and their crystallographic templates were less than 2 Å, suggesting that the modeling protocols yielded high-quality results. The models were compiled into an online resource, the Hemagglutinin Structure Prediction (HASP) server. The HASP server was designed as a scientific tool for researchers to visualize hemagglutinin protein sequences of interest in a three-dimensional context. With a built-in molecular viewer, hemagglutinin models can be compared side-by-side and navigated by a corresponding sequence alignment. The models and alignments can be downloaded for offline use and further analysis.Entities:
Mesh:
Substances:
Year: 2013 PMID: 23777206 PMCID: PMC3693987 DOI: 10.1186/1471-2105-14-197
Source DB: PubMed Journal: BMC Bioinformatics ISSN: 1471-2105 Impact factor: 3.169
HA crystal structures and templates selected for HASP server
| H1 | 1RD8, 1RU7, 1RUY, 1RUZ, |
| H2 | 2WR7, 2WRB, 2WRC, 2WRD, 2WRE, 2WRF, |
| H3 | 1EO8, 1HA0, 1HGD, 1HGE, 1HGF, 1HGG, 1HGH, 1HGI, 1HGJ, 1HTM, 1KEN, 1MQL, |
| H5 | |
| H7 | |
| H9 |
aPDB IDs in bold were selected as template structures for comparative modeling. The structures and the corresponding publications can be found online in the Research Collaboratory for Structural Biosciences PDB: http://www.rcsb.org/pdb.
Figure 1Rotamer recovery rate, sequence identity, and C-alpha RMSD for model-crystal structure pairs. The amino acid sequences for each crystal structure used as templates for the HASP server were modeled using the HASP server and the resultant models compared to the crystal structure. The identities are the number of identical amino acids between the query, template pair. The rotamer recovery rate is the ratio of residues with correctly modeled side-chains to all residues. The C-alpha RMSDs were calculated between the template structure and crystal structure of the sequence being modeled over all C-alpha atoms included in the model. Rotamer recovery versus identities can be fit by linear regression to the equation, rotamer recovery rate = 0.32 * identities + 0.41, with an R2 of 0.88. C-alpha RMSD versus rotamer recovery can be fit by linear regression to the equation, RMSD = −2.52 * identities + 3.16, with an R2 of 0.62. The PDB codes of the modeled sequence and template pairs are: A. 1JSD,3HTT; B. 1JSM,3GBM; C. 1MQM,2VIU; D. 1RV0,3GBN; E. 1RVX,3GBN; F. 1TI8,2VIU; G. 2VIU,1MQM; H. 3GBM,1JSM; I. 3GBN,1RV0; J. 3HTT,3GBN; K. 3KU3,3GBM; L. 3LZG,3GBN.
Summary of rotamer recovery for template sequences modeled using the HASP server
| 295 | 147 | 0.50 ± 0.11 | |
| 379 | 152 | 0.40 ± 0.13 | |
| 207 | 185 | 0.89 ± 0.11 | |
| 154 | 116 | 0.75 ± 0.12 | |
| 344 | 210 | 0.62 ± 0.10 | |
| 355 | 124 | 0.35 ± 0.13 | |
| 446 | 333 | 0.75 ± 0.08 | |
| 96 | 35 | 0.37 ± 0.18 | |
| 491 | 223 | 0.45 ± 0.11 | |
| 193 | 77 | 0.41 ± 0.12 | |
| 233 | 86 | 0.37 ± 0.11 | |
| 408 | 281 | 0.69 ± 0.14 | |
| 398 | 293 | 0.73 ± 0.11 | |
| 320 | 270 | 0.84 ± 0.07 | |
| 109 | 105 | 0.96 ± 0.09 | |
| 244 | 231 | 0.95 ± 0.06 |
aThe standard deviation (std) of the recovery rate was calculated over the set of individual recovery rates for each template sequence modeled (n = 12).
Figure 2Searching for hemagglutinin sequences. (A) Opening view in the Search tab. The H5 and N1 subtypes have been selected and the Toggle Map Search button is clicked. (B) The map selection viewer. Vietnam has been selected from the map and “2003 OR 2004” has been added to the search box. The Go button is clicked.
Figure 3Search results. A/Viet Nam/1203/2004 (AY818135) and A/chicken/Viet Nam/Ncvd8/2003 (EF541407) are selected and the Update View button is pressed. The full query is shown above the Select All button: TYPE = [H5] AND SUBTYPE = [*N1] AND KEYWORD = 2003 OR 2004 loc:[VN].
Figure 4The tab. Both sequences in the alignment are selected, activating the side-by-side view mode. The Single Chain button has been pressed to clarify the view. Position 145 is selected, centering the H5 models in the viewer and highlighting that position in cyan and its neighboring positions in yellow. The residue and its neighbors are displayed as sticks on top of the ribbon diagram. The Download button has been pressed, showing the drop-down menu in which Models and Sequences have been selected for download.