Literature DB >> 17073083

Genomic signatures of human versus avian influenza A viruses.

Guang-Wu Chen1, Shih-Cheng Chang, Chee-keng Mok, Yu-Luan Lo, Yu-Nong Kung, Ji-Hung Huang, Yun-Han Shih, Ji-Yi Wang, Chiayn Chiang, Chi-Jene Chen, Shin-Ru Shih.   

Abstract

Position-specific entropy profiles created from scanning 306 human and 95 avian influenza A viral genomes showed that 228 of 4591 amino acid residues yielded significant differences between these 2 viruses. We subsequently used 15,785 protein sequences from the National Center for Biotechnology Information (NCBI) to assess the robustness of these signatures and obtained 52 "species-associated" positions. Specific mutations on those points may enable an avian influenza virus to become a human virus. Many of these signatures are found in NP, PA, and PB2 genes (viral ribonucleoproteins [RNPs]) and are mostly located in the functional domains related to RNP-RNP interactions that are important for viral replication. Upon inspecting 21 human-isolated avian influenza viral genomes from NCBI, we found 19 that exhibited > or =1 species-associated residue changes; 7 of them contained > or =2 substitutions. Histograms based on pairwise sequence comparison showed that NP disjointed most between human and avian influenza viruses, followed by PA and PB2.

Entities:  

Mesh:

Substances:

Year:  2006        PMID: 17073083      PMCID: PMC3294750          DOI: 10.3201/eid1209.060276

Source DB:  PubMed          Journal:  Emerg Infect Dis        ISSN: 1080-6040            Impact factor:   6.883


Pandemic influenza A virus infections have occurred 3 times during the past century; the 1957 (H2N2) and 1968 (H3N2) pandemic strains emerged from a reassortment of human and avian viruses (). Recently, all 8 genome segments from the 1918 (H1N1) influenza A virus were completely sequenced. The results indicate that the 1918 pandemic virus may not have emerged by a reassortment of avian and human virus as did the 2 other pandemic strains. Although the 1918 H1N1 is not considered an avian virus, it is the most avianlike of all mammalian influenza viruses (,). The recent circulation of highly pathogenic avian H5N1 viruses in Asia from 2003 to 2006 has caused >90 human deaths and has raised concern about a new pandemic (). Therefore, we need to understand what genetic variations could render avian influenza virus capable of becoming a pandemic strain. Genomewide comparison of human versus avian influenza A viruses would show the evolutionary similarities and differences between them and thus provide information for studying the mechanism of influenza viral infection and replication in different host species. Although many research efforts have focused on the molecular evolution of specific genes of influenza viruses, comprehensive comparisons among the nucleotide sequences of all 8 genomic segments and among the 11 encoded protein sequences have not been extensively reported. In this study, we used several computational approaches for finding specific genetic signatures characteristic of human and avian influenza A viral genomes. We subsequently validated the robustness of those signatures with human and avian protein sequences downloaded from Influenza Virus Resources at the National Center for Biotechnology Information (NCBI) (http://www.ncbi.nlm.nih.gov/genomes/FLU/FLU.html).

Materials and Methods

Clinical Isolates

Throat swabs from patients with influenzalike syndromes were collected from the Clinical Virology Laboratory, Chang Gung Memorial Hospital. The specimens were inoculated in MDCK cells. Typing for influenza A virus was then performed with immunofluorescent assay by type-specific monoclonal antibody (Dako, Cambridgeshire, UK). Subtyping was conducted by reverse transcription (RT)–PCR with subtype-specific primers.

Sequence Analysis

The RT-PCR product was purified by using the QIAquick Gel Extraction Kit (Qiagen, Valencia, CA, USA). The nucleotide sequence was determined with an automated DNA sequencer. Sequence editing and processing were performed with Lasergene, version 3.18 (DNASTAR, Madison, WI, USA). Multiple sequence alignment was performed with ClustalW version 1.83 (ftp://ftp.ebi.ac.uk/pub/software/unix/clustalw). Global sequence comparison that yielded pairwise sequence identities used in histogram analysis was done with the program Needle in the EMBOSS package (). Amino acid sequences were translated from coding sequences and aligned by BioEdit (). An entropy value was defined at an aligned amino acid position according to the formula ΣP(P), in which i is the observed probability for each of the 20 amino acids (aa) (). A graphic tool was developed in Java for displaying the entropy plot used in this work. All amino acid numberings are based on influenza virus A/Puerto Rico/8/1934 (PR8).

Sequences Used in Study

To show the host-associated amino acid signatures, we retrieved full genome sequences (as of August 22, 2005) from the genome browser at Influenza Sequence Database (ISD) (). To differentiate between avian and human influenza viruses, we excluded human-isolated avian influenza viruses from the human dataset and examined those sequences separately. Altogether, we had 95 avian and 306 human influenza viral genomes, henceforth termed "primary dataset." All 11 viral proteins encoded by the 8 genomic RNA segments were compared: PB2, PB1, PB1-F2, PA, HA, NP, NA, M1, M2, NS1, and NS2. Avian influenza viruses from human influenza patients were separately retrieved from NCBI as well as from ISD. Altogether, we had 417 protein sequences from 60 avian influenza strains, in which 21 strains contain sequences (full or nearly full length) from all 8 genomic RNA segments. For validating the signatures obtained from analyzing the primary dataset, we further retrieved 15,785 human or avian influenza A viral protein sequences from NCBI's Influenza Virus Resources. Details for the sequences used can be found in Appendix, Supporting Materials and Methods, as well as in Table A1 and Table A2. Eleven Taiwanese genomes produced in this work have been deposited in GenBank with accession numbers DQ415283 through DQ415370.
Table A1

Listing of 401 genomes used in this study. All accessions are according to GenBank, except for A/Puerto Rico/8/34(H1N1), which are from Influenza Sequence Database (ISD). Full table available at www.cdc.gov/eid-static/spreadsheets/06-0276-TA1.xlsx.

StrainSubtypeHostPB2PB1PAHANPNAMNS
A/BAR-HEADED GOOSE/QINGHAI/5/05H5N1AvianDQ095757DQ095737DQ095717DQ095617DQ095677DQ095657DQ095637DQ095697
A/BAR-HEADED GOOSE/QINGHAI/59/05H5N1AvianDQ095752DQ095732DQ095712DQ095612DQ095672DQ095652DQ095632DQ095692
A/BAR-HEADED GOOSE/QINGHAI/60/05H5N1AvianDQ095755DQ095735DQ095715DQ095615DQ095675DQ095655DQ095635DQ095695
A/BAR-HEADED GOOSE/QINGHAI/61/05H5N1AvianDQ095758DQ095738DQ095718DQ095618DQ095678DQ095658DQ095638DQ095698
A/BAR-HEADED GOOSE/QINGHAI/62/05H5N1AvianDQ095760DQ095740DQ095720DQ095620DQ095680DQ095660DQ095640DQ095700
A/BAR-HEADED GOOSE/QINGHAI/65/05H5N1AvianDQ095762DQ095742DQ095722DQ095622DQ095682DQ095662DQ095642DQ095702
A/BAR-HEADED GOOSE/QINGHAI/67/05H5N1AvianDQ095763DQ095743DQ095723DQ095623DQ095683DQ095663DQ095643DQ095703
A/BAR-HEADED GOOSE/QINGHAI/68/05H5N1AvianDQ095753DQ095733DQ095713DQ095613DQ095673DQ095653DQ095633DQ095693
A/BAR-HEADED GOOSE/QINGHAI/75/05H5N1AvianDQ095759DQ095739DQ095719DQ095619DQ095679DQ095659DQ095639DQ095699
A/BIRD/THAILAND/3.1/2004H5N1AvianAY651715AY651661AY651607AY651330AY651495AY651441AY651384AY651550
A/BROWN-HEADED GULL/QINGHAI/3/05H5N1AvianDQ095756DQ095736DQ095716DQ095616DQ095676DQ095656DQ095636DQ095696
A/CHICKEN/BEIJING/1/94H9N2AvianAF156438AF156423AF156452AF156380AF156409AF156398AF156466AF156480
A/CHICKEN/BEIJING/8/98H9N2AvianAF508649AF508627AF508671AF508562AF508605AF508583AF508693AF508714
A/CHICKEN/BRITISH COLUMBIA/04H7N3AvianAY616766AY616765AY616764AY611524AY611527AY611526AY611525AY611528
A/CHICKEN/CALIFORNIA/139/01H6N2AvianAF457705AF457706AF457707AF457713AF474070AF457711AF457712AF457708
A/CHICKEN/CALIFORNIA/431/00H6N2AvianAF457697AF457698AF457699AF457704AF457701AF457702AF457703AF457700
A/CHICKEN/CALIFORNIA/465/00H6N2AvianAF457689AF457690AF457691AF457696AF457693AF457694AF457695AF457692
A/CHICKEN/CALIFORNIA/6643/01H6N2AvianAF457681AF457682AF457683AF457688AF457685AF457686AF457687AF457684
A/CHICKEN/CALIFORNIA/905/01H6N2AvianAF457672AF457673AF457674AF457679AF457676AF457677AF457678AF457675
A/CHICKEN/GERMANY/R28/03H7N7AvianAJ620347AJ620348AJ619677AJ620350AJ620352AJ620349AJ619676AJ619678
A/CHICKEN/GUANGDONG/10/00H9N2AvianAF508650AF508628AF508672AF508563AF508606AF508584AF508694AF508715
A/CHICKEN/GUANGDONG/11/97H9N2AvianAF508651AF508629AF508673AF508564AF508607AF508585AF508695AF508716
A/CHICKEN/GUANGDONG/174/04H5N1AvianAY609309AY609310AY609311AY609312AY609313AY609314AY609315AY609316
A/CHICKEN/GUANGDONG/178/04H5N1AvianAY737293AY737294AY737295AY737296AY737297AY737299AY737298AY737300
A/CHICKEN/GUANGDONG/191/04H5N1AvianAY737286AY737287AY737288AY737289AY737290AY737291AY737292AY737285
A/CHICKEN/HONG KONG/220/97H5N1AvianAF046086AF046085AF046087AF046080AF046084AF046081AF046082AF046083
A/CHICKEN/HONG KONG/728/97H5N1AvianAF098579AF098592AF098606AF046099AF098618AF098548AF098562AF098571
A/CHICKEN/HONG KONG/739/94H9N2AvianAF156436AF156422AF156450AF156379AF156408AF156397AF156464AF156478
A/CHICKEN/HONG KONG/FY150/01H5N1AvianAY221587AY221578AY221569AY221524AF509120AF509095AF509043AY221560
A/CHICKEN/HONG KONG/NT873.3/01H5N1AvianAY221585AY221576AY221567AY221522AY221549AY221540AY221531AY221558
A/CHICKEN/HONG KONG/YU562/01H5N1AvianAY221592AY221583AY221574AY221529AF509118AY221547AF509041AF509067
A/CHICKEN/HONG KONG/YU822.2/01H5N1AvianAY221591AY221582AY221573AY221528AY221555AY221546AY221537AY221564
Table A2

Listing of 60 human-isolated avian influenza viruses used in this study, with the first 29 strains contain at least one accession per genomic segment. All accessions here are according to GenBank, except those ones begin with 'ISD', which are from Influenza Sequence Database (ISD). Cells with 'n/a' in PB1-F2 column indicate that the PB1 RNA sequence did not contain the PB1-F2 ORF, while 'truncated' represent early-terminated PB1-F2 with a length less than 87-aa. Exclusion of those PB1-F2 leaves us 21 genomes for inspecting the species-associated mutations as described in the text.

Strain NamePB2PB1PB1-F2PAHANPNAM1M2NS1NS2
A/HongKong/156/97(H5N1)AF036363AF036362AF036362AF084267AF028709AF028710AF036357AF036358AF036358AF036360AF036360
A/HongKong/481/97(H5N1)AF115290AF258818AF258818AF115294AF046096AJ289873AF084271AF115286AF115286AF115288AF115288
A/HongKong/482/97(H5N1)AF258838AF258819AF084264AF084268AF046098AF255745AF084272AF084282AF084282AF084285AF084285
A/HongKong/483/97(H5N1)AF258839AF258820AF084265AF084269AF046097AF084277AF084273AF255367AF255367AF084286AF084286
A/HongKong/485/97(H5N1)AF084263AF084266truncatedAF084270AF102681AF084278AF084274AF084284AF084284AF084287AF084287
A/HongKong/486/97(H5N1)AF115291AF115293AF115293AF115295AF102671AF115285AF084275AF255368AF255368AF256181AF256181
A/HongKong/488/97(H5N1)AF258848AF258829n/aAF257204AF102672AF255756AF102657AF255377AF255378AF256190AF256190
A/HongKong/491/97(H5N1)AF258849AF258830n/aAF257205AF102677AF255758AF102665AF255379AF255380AF256191AF256191
A/HongKong/503/97(H5N1)AF258850AF258831n/aAF257206AF102679AF255760AF102666AF255381AF255381AF256192AF256192
A/HongKong/507/97(H5N1)
AF258851
AF258832
n/a
AF257207
AF102675
AF255762
AF102659
AF255382
AF255382
AF256193
AF256193
A/HongKong/514/97(H5N1)AF258852AF258833n/aAF257208AF102682AF255764AF102669AF255383AF255383AF256184AF256184
A/HongKong/516/97(H5N1)AF258853AF258834n/aAF257209AF102673AF255766AF102660AF255384AF255384AF256194AF256194
A/HongKong/532/97(H5N1)AF258843AF258824AF258824AF257199AF102680AF255750AF102667AF255371AF255371AF256185AF256185
A/HongKong/538/97(H5N1)AF258844AF258825AF258825AF257200AF102674AF255751AF102662AF255372AF255372AF256186AF256186
A/HongKong/542/97(H5N1)AF258845AF258826AF258826AF257201AF102678AF255752AF102670AF255373AF255373AF256187AF256187
A/HongKong/97/98(H5N1)AF258846AF258827AF258827AF257202AF102676AF255753AF102661AF255374AF255374AF256188AF256188
A/HongKong/212/03(H5N1)AY576380AY576392AY576392AY576404AY575869AY575905AY575881AY575893AY575893AY576368AY576368
A/HongKong/213/2003(H5N1)AY576381AB212052AY576393AB212053AB212054AB212055AB212056AB212057AB212057AY576369AY576369
A/Thailand/16/2004(H5N1)ISDN40383ISDN40859ISDN40859ISDN40940ISDN40341ISDN40086ISDN48790ISDN45755ISDN45755ISDN40040ISDN40040
A/Thailand/SP83/2004(H5N1)
ISDN49457
ISDN40931
ISDN40931
ISDN121933
ISDN40917
ISDN41067
ISDN48792
ISDN111182
ISDN111182
ISDN41028
ISDN41028
A/Vietnam/1194/2004(H5N1)AY651718AY651664AY651664AY651610AY651333AY651498ISDN38703ISDN39957ISDN39957AY651552AY651552
A/Vietnam/1196/04(H5N1)AY526752AY526751AY526751AY526750AY526745AY526749AY526746AY526748AY526748AY526747AY526747
A/Vietnam/1203/2004(H5N1)AY651719AY818129AY651665AY818132ISDN38687AY818138AY651447AY651388AY651388AY651553AY651553
A/Vietnam/3046/2004(H5N1)AY651720AY651666AY651666AY651613AY651335AY651500AY651446AY651389AY651389AY651554AY651554
A/Vietnam/3062/2004(H5N1)AY651721AY651667AY651667AY651612AY651336AY651501AY651448AY651390AY651390AY651555AY651555
A/Netherlands/219/03(H7N7)AAR04358AAR05983AY340083AAR04363AAR02640AAR04370AAR11367AAR11371AY340089AAR04367AY342422
A/Guangzhou/333/99(H9N2)AY043030AY043029truncatedAY043028AY043019AY043026AY043024AY043025AY043025AY043027AY043027
A/HongKong/1073/99(H9N2)AF258835AF258816AF258816AF257191AJ404626AJ289871AJ404629AF255363AF255363AJ278649AJ278649
A/HongKong/1074/99(H9N2)AF258836AF258817AF258817AF257192AJ404627AJ289872AJ404628AF255364AF255364AF256177AF256177
A/England/268/96(H7N7)
 
 
 
 
AF028020
 
 
 
 
 
 
A/Shantou/239/98(H9N2)    AY043015 AY043021    
A/Shaoguan/408/98(H9N2)    AY043017 AY043022    
A/Shaoguan/447/98(H9N2)    AY043018 AY043023    
A/unknown/149717-12/2002(H7N2)       DQ107480DQ107480  
A/Netherlands/124/03(H7N7)AAR04355AAR05980AY340080AAR04360  AAR11364AAR11368AY340086  
A/Netherlands/126/03(H7N7)AAR04356AAR05981AY340081AAR04361  AAR11363AAR11369AY340087  
A/Netherlands/127/03(H7N7)AAR04357AAR05982AY340082AAR04362AAR02636  AAR11370AY340088AAR04366AY342421
A/Netherlands/33/03(H7N7) AAR05984AY340084AAR04364AAR02638AAR04371AAR11366AAR11372AY340090AAR04368AY342423
A/Hanoi/03/2004(H5N1)    AJ715872AJ715873     
A/Hatay/2004(H5N1)
 
 
 
 
AJ867074
AJ867076
AJ867075
AM040045
AM040045
AM040046
AM040046
A/Prachinburi/6231/2004(H5N1)    ISDN110940 ISDN110939    
A/Thailand/1-KAN-1/2004(H5N1)    AY555150 AY555151    
A/Thailand/2-SP-33/2004(H5N1)    AY555153 AY555152    
A/Thailand/Chaiyaphum/622/2004(H5N1)    ISDN49460 ISDN48793ISDN111184ISDN111184  
A/Thailand/EKA2NF/2004(H5N1)      AY535029    
A/Thailand/Kamphaengphet-Nontaburi/04(H5N1)    AY786078 AY786079    
A/Thailand/Kan353/2004(H5N1)    ISDN40918 ISDN48791ISDN111183ISDN111183  
A/Thailand/Prachinburi/6231/2004(H5N1)       ISDN111185ISDN111185  
A/Thailand/LFPN-2004/2004(H5N1)    AY679514 AY679513    
A/Vietnam/1194/2004(H5N1)
 
 
 
 
ISDN38686
 
AY651445
AY651387
AY651387
 
 
A/Vietnam/1204/2004(H5N1)ISDN40380ISDN40843ISDN40843ISDN121932ISDN38688    ISDN40017ISDN40017
A/Vietnam/3212/2004(H5N1)    ISDN40278      
A/Vietnam/DN-33/2004(H5N1)    AY720950 AY720948  AY720949AY720949
A/Vietnam/JP178/2004(H5N1)    ISDN69608 ISDN69610    
A/Vietnam/HN/2004(H5N1)AY720954AY720955n/aAY720952 AY720953 AY720951AY720951  
A/Cambodia/JP52a/2005(H5N1)    ISDN121986 ISDN122818    
A/Hanoi/30408/2005(H5N1)    ISDN129400      
A/Vietnam/HN30408/2005(H5N1)    ISDN119678 ISDN119679    
A/Vietnam/JP14/2005(H5N1)    ISDN117778 ISDN117783    
A/Vietnam/JP4207/2005(H5N1)    ISDN117777 ISDN117782    
A/Vietnam/JPHN30321/2005(H5N1)    ISDN118371      

Results

Differing Amino Acid Residues

Using previously described methods (), we separately calculated an entropy value for every aligned amino acid position for 95 avian influenza viruses and 306 human influenza viruses. Those amino acid residues with an entropy value between 0 and –0.4 for both the human and avian strains were identified as most highly conserved. We chose this entropy threshold on the basis of the entropy value –0.379, calculated at position 627 of PB2 for the 95 avian viruses. This widely reported, species-associated residue is highly conserved; it has E (Glu) in 83 and K (Lys) in 12 avian isolates and Lys in all 306 human isolates. We then selected those conserved positions with distinct amino acid residues between human and avian influenza viruses as potential host-associated signatures. An entropy plot for identifying such signature residues for avian versus human influenza virus NP segments is shown in Figure panel A. In each aligned position, we placed an avian consensus residue on top and a human consensus at the bottom. For example, the entropy value is zero at amino acid position 283 for both avian and human strains, in which all 95 avian influenza viruses contain L (Leu), whereas all 306 human influenza viruses contain P (Pro). The other 2 residues with zero entropy value in avian and human viruses are located at position 55 of PA, in which we have D (Asp) in avian viruses and N (Asn) in human viruses, and position 121 of M1, in which we have T (Thr) in avian and A (Ala) in human viruses. Entropy plots for all 11 influenza viral proteins can be found in Figure A1. Figure panel B shows a genomewide view of the entropy plots for 11 influenza A viral proteins. The amino acid sequences of hemagglutinin (HA), with an average entropy value of –0.524 within avian viruses and –0.158 within human viruses, exhibit much more diversity than other open reading frames (ORFs). PB2, PB1, PA, NP, and M1, on the other hand, are more conserved (i.e., they have less negative entropy values). A) Entropy plot for avian versus human influenza viruses for NP amino acid residues. In each aligned position, we have a consensus residue for 95 avian strains displayed on top and a consensus residue for 306 human strains at the bottom. Completely conserved amino acid positions are filled with white; less conserved amino acids are filled in various gray shadings. Positions in which 1 single residue dominates >90%, <90% but >75%, and <75% are labeled with red, yellow, and green letters, respectively. Yellow rectangles indicate that both human and avian viruses are completely conserved to the same residue; magenta rectangles indicate that avian and human viruses are each completely conserved to a different residue. B) Entropy plots for the entire influenza A viral genome. Each lane displays entropy value distributions of aligned protein sequences for 1 of the 11 viral proteins; the upper half represents 95 avian strains, and the bottom half represents 306 human strains. (PB1-F2 contains fewer strains, as described in Discussion.) Positions completely conserved to a single residue are shown in a white band, while less conserved ones are shown in various gray shadings. The average entropy for the entire segment is shown to the right of these lanes. Entropy values are zero when residues are completely conserved; more negative values indicate more diversity. Alignment size for each protein from top to bottom is 759, 757, 90, 716, 591, 498, 480, 252, 97, 230, and 121. In addition to the previously mentioned 3 positions with distinct amino acid residues between avian and human strains, we found 225 additional positions with nearly distinct amino acid residues, with their computed entropy values less negative than –0.4 in both the 306 human and 95 avian strains that we analyzed. To assess the robustness of those 228 residues used in differentiating human from avian influenza viruses, we further examined 15,785 influenza A protein sequences from NCBI. After validation, 52 positions still showed an entropy value less negative than –0.4 and conserved to distinct amino acid residues between human and avian viruses (Table 1). From this entropy analysis, we identified an additional 51 aa positions that may be as important as the well-known position 627 of PB2. We designated these 52 positions as "species-associated" signatures. Among 11 ORFs, NP contains the highest number of such signatures (15 positions), followed by PA (10 positions), PB2 (8 positions), PB1-F2 (5 positions), M2 (4 positions), M1 (3 positions), PB1 (2 positions), HA (2 positions), NS2 (2 positions), and NS1 (1 position). No signature was found in the NA gene. We also summarized the related functions of those species-associated signatures in Table 1. The complete results of genome scanning and validation can be found in Table A3 and Table A4.
Table 1

Validated amino acid signatures separating avian influenza viruses from human influenza viruses*

GenePositionAvian residuesHuman residuesAssociated functional domains
PB244A(208),S(7)S(831),A(10),L(2)PB1–1, NP-1 (9), MLS (10)
199A(210),S(5)S(842),A(3)NP-1 (9)
271T(210),A(3),I(1),M(1)A(836),T(6),S(1)Cap-N (11)
475L(214),M(1)M(839),L(3)NLS (12)
588A(203),T(6),V(6)I(835),V(3),A(2)PB1–2, NP-2 (9)
613V(212),A(3)T(816),I(16),A(8),V(1)PB1–2, NP-2 (9)
627E(196),K(19)K(838),R(2),E(1)PB1–2, NP-2 (9)
674A(204),S(6),T(2),G(2),E(1)T(836),A(2),I(2),P(1)PB1–2, NP-2 (9)
PB1327R(147),K(3)K(766),R(66)cRNA (13)
336V(142),I(8)I(773),V(59)cRNA (13)
PB1-F273K(397),R(6),I(1)R(594),K(87),S(1)ANT3, VDAC1 (14), mitochondrial localization (15), predicted amphipathic helix (16)
76V(401),A(3)A(625),V(57)ANT3, VADC1 (14), predicted amphipathic helix (16)
79R(369),Q(34),L(1)Q(607),R(75)ANT3, VADC1 (14), predicted amphipathic helix (16)
82L(382),S(22)S(596),L(86)ANT3, VADC1 (14), predicted amphipathic helix (16)
87E(389),G(14),K(1)G(637),E(45)ANT3, VADC1 (14)
PA28P(213),S(1)L(831),P(9),R(2)Proteolysis (17)
55D(214)N(836),D(5)Proteolysis (17)
57R(210),Q(4)Q(829),R(6),L(4),K(2)Proteolysis (17)
225S(213),C(1)C(829),S(10)Proteolysis (17), NLSII (18)
268L(214)I(827),L(11), P(1)
356K(212),X(1),R(1)R(827),K(11)
382E(208),D(5),V(1)D(824),E(11),V(2),N(1)
404A(214)S(828),A(9),P(1)
409S(189),N(24),I(1)N(830),S(7),I(1)
552T(213),N(1)S(835),T(1),I(1)
HA237N(582),R(49),D(2),H(1),S(1)R(1209),N(12),S(2),D(1),K(1)
389D(659),N(20),G(1),Y(1)N(819),D(121)
NP16G(356),S(9),D(6),T(2)D(646),G(7)RNA binding (19), BAT1/UAP56 (20), MxA (21), PB2–1 (22)
33V(355),I(18)I(638),V(15)RNA binding (19), MxA (21), PB2–1 (22)
61I(366),M(6),V(1)L(642),I(8)RNA binding (19), MxA (21), PB2–1 (22)
100R(360),K(11),V(2)V(619),I(32),A(1),M(1)RNA binding (19), MxA (21), PB2–1 (22)
109I(359),V(10),M(2),T(2)V(614),I(34),T(3),A(2)RNA binding (19), MxA (21), PB2–1 (22)
214R(352),K(20),L(1)K(640),R(10)NLS (23), CRM1 (24), NP-1 (25)
283L(372),P(1)P(643),L(7)NP-1 (25), PB2–2 (22)
293R(371),K(2)K(622),R(28)NP-1 (25), PB2–2 (22)
305R(369),K(4)K(636),R(14)NP-1 (25), PB2–2 (22)
313F(371),I(1),L(1)Y(642),F(8)NP-1 (25), PB2–2 (22)
357Q(368),K(4),T(1)K(644),R(8),Q(1)NAS (26), NP-1 (25), PB2–3 (22)
372E(357),D(15),K(1)D(630),E(23)NAS (26), NP-2 (25), PB2–3 (22)
422R(373)K(630),R(23)CTL epitope (27), NP-2 (25), PB2–3 (22)
442T(372),A(1)A(629),T(23),R(1)NP-2 (25), PB2–3 (22)
455D(373)E(630),D(22),T(1)NP-2 (25), PB2–3 (22)
M1115V(856),I(2),L(1),G(1)I(981),V(9)
121T(840),A(19),P(1)A(988),T(2)
137T(859),A(1),P(1)A(974),T(12)
M211T(434),I(11),S(2)I(911),T(44)Host restriction specificities (28), ectodomain (29)
20S(471),N(13)N(926),S(29)Host restriction specificities (28). ectodomain (29)
57Y(481),C(1),H(1)H(913),Y(33),R(2),Q(1)CRAC (30), endodomain (29)
86V(378)A(924),V(10),T(4),D(1)Endodomain (29)
NS1227E(692),G(9),K(1),S(1)R(897),G(5),K(1),E(1)
NS270S(453),G(21),D(1)G(903),S(2)M1, NEP dimerization domain (31)
107L(468),S(2),F(1)F(777),L(16),S(1)M1, NEP dimerization domain (31)

*Numbers in parentheses in residue columns are the number of sequences yielding the specific amino acid residue; bold indicates dominant amino acid residue type.

Table A3

Genome-scanning for 228 amino acid 'signatures' (those ones shown in bold face, either 'Distinct' or 'Nearly Distinct') out of 4,591 aligned amino acid positions. Con: consensus residue; Ent: entropy value; Same: all avian and human strains conserve to the same residue; Nearly Identical: both avian and human strains have entropy values less negative than -0.4, and have the same residue; Distinct: both avian and human strains contain zero entropy yet conserve to the different residue; Nearly Distinct: both avian and human strains contain an entropy value less negative than -0.4, and conserve to different residue. The last column shows positions based on PR8. Only HA and NA have different numberings comparing with the first column, due to excessive shifting of amino residues from HA and NA genetic diversity. Full table available at www.cdc.gov/eid-static/spreadsheets/06-0276-TA3.xlsx.

Scanning for amino acid 'signatures' for influenza A virus PB2 protein
 
PosAvian
Human
CommentsPR8
ConEntResiduesConEntResidues
1M-0.500M(76),-(19),M-0.055M(303),-(3),
2E-0.293E(88),K(1),-(6),E-0.061E(303),V(1),-(2),Nearly Identical
3R-0.175R(91),-(4),R-0.022R(305),T(1),Nearly Identical
4I-0.140I(92),-(3),I-0.022I(305),L(1),Nearly Identical
5K-0.140K(92),-(3),K-0.039R(2),K(304),Nearly Identical
6E-0.140E(92),-(3),E0.000E(306),Nearly Identical
7L-0.160L(92),F(1),-(2),L0.000L(306),Nearly Identical
8R-0.160R(92),W(1),-(2),R0.000R(306),Nearly Identical
9D-0.521N(1),D(83),E(7),Y(2),-(2),N-0.619N(227),D(1),S(2),T(76),
10L-0.160I(1),L(92),-(2),L-0.022I(1),L(305),Nearly Identical
11M-0.117I(1),M(93),-(1),M0.000M(306),Nearly Identical
12S0.000S(95),S-0.022L(1),S(305),Nearly Identical
13Q0.000Q(95),Q0.000Q(306),Same
14S0.000S(95),S-0.022F(1),S(305),Nearly Identical
15R0.000R(95),R0.000R(306),Same
16T-0.058S(1),T(94),T0.000T(306),Nearly Identical
17R0.000R(95),R0.000R(306),Same
18E0.000E(95),E0.000E(306),Same
19I0.000I(95),I0.000I(306),Same
20L0.000L(95),L-0.022L(305),V(1),Nearly Identical
21T0.000T(95),T0.000T(306),Same
22K0.000K(95),K-0.022N(1),K(305),Nearly Identical
23T0.000T(95),T-0.022P(1),T(305),Nearly Identical
24T0.000T(95),T0.000T(306),Same
25V0.000V(95),V0.000V(306),Same
Table A4

Amino acid 'signatures' validation. Only positions with both newly computed entropy values less or equal to -0.400 are considered 'validated'. This reduces 228 signatures to a count of 52 (the ones shown in bold face) as reported in the manuscript. Cnt: total number of avian or human residues at this position; Ent: computed entropy value; PR8: position numbering based on PR8 (only HA and NA have different numbering here). Full table available at www.cdc.gov/eid-static/spreadsheets/06-0276-TA4.xlsx.

GenePosAvian influenza viruses
Human influenza viruses
Validated?PR8
CntEntResiduesCntEntResidues
PB244215-0.144 A(208),S(7),843-0.081 A(10),L(2),S(831),Yes 
 67215-0.174 I(206),V(9),843-0.538 I(193),V(650),  
 81215-0.196 A(2),I(7),T(206),843-0.537 I(2),L(4),M(686),T(4),V(147),  
 82215-0.164 R(1),N(209),K(1),S(2),T(2),843-0.608 N(180),C(14),S(648),X(1),  
 1202150.000 E(215),843-0.575 N(1),D(628),E(214),  
 199215-0.110 A(210),S(5),845-0.024 A(3),S(842),Yes 
 2272150.000 V(215),845-0.697 I(586),M(19),V(240),  
 271215-0.133 A(3),I(1),M(1),T(210),843-0.051 A(836),S(1),T(6),Yes 
 382215-0.110 I(210),V(5),842-0.527 I(185),V(657),  
 453215-0.164 Q(2),H(1),L(2),P(209),S(1),842-0.497 R(1),H(691),L(1),P(147),S(2),  
 456215-0.219 N(205),D(6),S(4),842-0.623 N(247),D(1),C(1),S(593),  
 461215-0.159 I(207),V(8),842-0.638 I(283),V(559),  
 463215-0.247 I(203),L(1),M(1),V(10),842-0.529 I(181),M(1),V(660),  
 475215-0.030 L(214),M(1),842-0.024 L(3),M(839),Yes 
 478215-0.484 I(30),L(1),M(2),V(182),842-0.541 I(656),L(2),V(184),  
 526215-0.053 R(2),K(213),841-0.577 R(619),K(222),  
 559215-0.255 I(5),M(2),T(204),V(4),841-0.694 A(547),N(1),I(2),T(287),V(4),  
 588215-0.254 A(203),T(6),V(6),841-0.050 A(2),I(835),V(3),X(1),Yes 
 613215-0.073 A(3),V(212),841-0.157 A(8),I(16),T(816),V(1),Yes 
 627215-0.299 E(196),K(19),841-0.026 R(2),E(1),K(838),Yes 
*Numbers in parentheses in residue columns are the number of sequences yielding the specific amino acid residue; bold indicates dominant amino acid residue type.

Amino Acid Signatures in Human Viruses

We examined how the amino acid sequences varied at those proposed signature positions for avian influenza viruses isolated from humans. At 9 of these 52 positions, residue changes were characteristic of human rather than avian viruses (Table 2). For example, 34 sequences (27 H5N1, 3 H9N2, and 4 H7N7) were available for inspection at position 199 of PB2 (data not shown). Aside from 10 sequences with gaps (sequences did not cover this position), 19 of the remaining 24 still have Ala, which is typical for avian viruses. Five of them (all H5N1), on the other hand, have this residue changed to Ser, which is mostly seen in human viruses. At the well-known position 627 of PB2, 5 sequences had gaps, 22 retained Glu (typical for avian virus), while the other 7 changed to Lys, which is typical for human virus. Among those 7 mutated sequences, 6 were from H5N1 human isolates (A/Hong Kong/483/1997, A/Hong Kong/485/1997, A/Vietnam/1194/2004, A/Vietnam/1203/2004, A/Vietnam/3062/2004, and A/Thailand/16/2004), and the other 1 was A/Netherlands/219/2003(H7N7), which was isolated from a fatal human case of pneumonia in the Netherlands ().
Table 2

Summary of host-associated amino acid signature changes

GenePositionResidue*H5N1H9N2H7N2H7N7
PB2199A(19)1531
S(5)5
271T(23)2021
A(1)1
627E(22)193
K(7)61
PB1-F273K(24)1725
R(2)2
79R(24)1725
Q(2)2
82L(21)192
S(5)5
PA409S(17)1232
N(7)7
M220S(34)3121
N(5)5
NS270S(26)2222
G(1)1

*Top half displays an avian-specific residue with the count in parentheses and distribution among subtypes, and the bottom half represents a human-specific residue.

*Top half displays an avian-specific residue with the count in parentheses and distribution among subtypes, and the bottom half represents a human-specific residue. To understand how mutations had accumulated within a specific virus, we summarized the amino acid changes for 21 of these avian viruses that contained full or nearly full-length sequences for each segment (Table 3). We found that 19 of 21 strains contained >1 species-associated amino acid change, and 7 of them contained >2 substitutions; A/Netherlands/219/2003(H7N7) had the highest count for mutation accumulation (3 positions). Among these 52 species-associated signatures, the mutation combinations at positions PB2 199 and PA 409 were most commonly seen in H5N1 human isolates from Hong Kong in 1997.
Table 3

Twenty-one avian influenza A viral genomes isolated from humans and their mutations found at 12 host-associated positions within each strain*

StrainSubtypePB2
PB1-F2
PA
M2
NS2
Mutations
1992716277379824092070
A/Hong Kong/156/1997H5N1 S TEKRL N SS2
A/Hong Kong/481/1997H5N1ATEKRL N SS1
A/Hong Kong/482/1997H5N1 S TEKRL N SS2
A/Hong Kong/483/1997H5N1AT K KRLSSS1
A/Hong Kong/485/1997H5N1AT K ###SSS1
A/Hong Kong/486/1997H5N1 S TEKRL N SS2
A/Hong Kong/532/1997H5N1ATEKRL N SS1
A/Hong Kong/538/1997H5N1 S TEKRL N SS2
A/Hong Kong/542/1997H5N1ATEKRL N SS1
A/Hong Kong/1997/1998H5N1 S TEKRLSSS1
A/Hong Kong/212/2003H5N1ATE R RLSSS1
A/Hong Kong/213/2003H5N1ATE R RLSSS1
A/Thailand/16/2004H5N1AT K K Q LSSS2
A/Thailand/SP83/2004H5N1ATEK Q LSSS1
A/Vietnam/1194/2004H5N1AT K KRLSSS1
A/Vietnam/1203/2004H5N1AT K KRLSSS1
A/Vietnam/3062/2004H5N1AT K KRLSSS1
A/Netherlands/219/2003H7N7AT K KR S S N S3
A/Guangzhou/333/1999H9N2A A E###SS G 2
A/Hong Kong/1073/1999H9N2ATEKRLSRS0
A/Hong Kong/1074/1999H9N2ATEKRLSSS0

*#indicates strains with PB1 RNA encoded into a truncated form of PB1-F2 of only 57 amino acids long. Boldface letters represent mutated (human-specific) residues; Roman (nonbold) letters are used for regular avian residue. Note that at position 20 of M2, A/Hong Kong/1073/99 had its residue changed from S to R, where R is still considered a mutation within avian species.

*#indicates strains with PB1 RNA encoded into a truncated form of PB1-F2 of only 57 amino acids long. Boldface letters represent mutated (human-specific) residues; Roman (nonbold) letters are used for regular avian residue. Note that at position 20 of M2, A/Hong Kong/1073/99 had its residue changed from S to R, where R is still considered a mutation within avian species.

RNA Segment 5

Our observation that NP contained the highest number (15 of 52) for species-associated amino acids suggested that NP might serve as a molecular target for differentiation between human and avian influenza A viruses. To indicate such host specificity, or the "genetic boundary" between these 2 viruses at the nucleotide level, we performed a pairwise sequence comparison for all 11 ORFs on our 401-genome primary dataset and produced histograms on their computed pairwise identities. In Figure A2, pairs with 2 sequences of the same host species (human to human, or avian to avian; termed homopairs) and pairs for sequences that cross host species (human to avian, or avian to human; termed heteropairs) are shown. HA and NA genes exhibited considerable sequence differences between strains, with identities as low as 47%. Also noted was a wide spectrum of percent identities (e.g., 55%–95% in the horizontal axis) containing few sequence pairs for these 2 genes. For both of these proteins, some strains from the same species can have identities as low as 50%. However, the ORF of another surface protein, M2 ion channel protein, is relatively conserved (>74% identity for viruses across species). The histograms for the polymerase genes (PB2, PB1, and PA), NP, and M1, on the other hand, are much less varied (mostly <20% variation). In particular, the NP gene was found to exhibit a fairly clear boundary between homopairs and heteropairs, at ≈86%.
Figure A2

Histograms on comparing 306 human versus 95 avian influenza A viruses, based on nucleotide pairwise sequence identities. Vertical axis shows the count for pairs of sequences with specific percent identity (rounded to integer). Red bars represent frequencies for 'homo' pairs – sequences of the same host species (human to human, or avian to avian); blue bars represent frequencies for 'hetero' pairs – pairs that cross host species (human to avian, or avian to human). Adobe Acrobat PDF available at http://wwwnc.cdc.gov/eid/pdfs/06-0276-FA2.pdf (6 pages).

Discussion

The glutamic acid residue at PB2 627, which is commonly seen in avian viruses, restricts viral growth in humans and monkeys, but a change to lysine restores virus replication in mammalian cells (). In this study we computed for every amino acid position (distributed in the 11 known influenza viral ORFs) an entropy value that represents how conserved an amino acid residue is at that given position. We found the entropy value –0.379 at 627 of PB2 and therefore used –0.4 as a threshold to discover other amino acid residues that might be potential determinants of host-cell tropism. Another 51 positions were found to be distinct or nearly distinct between human and avian viruses by this entropy threshold. Most of these (40 of 52) are located in viral ribonucleoproteins (RNPs) (PB2, PB1, PA, and NP), which are essential for viral replication. Taubenberger et al. reported 10 amino acid residues that distinguish human and avian influenza viral polymerases (). Six of them were also identified in this study. The entropy values of the 4 missing ones were also found close to the preset threshold (–0.4). For example, PB2 567 showed a human entropy of –0.039 and avian entropy of –0.490, PB1 375 with human entropy –0.165 and avian entropy –0.693, and PA 100 with human entropy –0.061 and avian entropy –0.406. All 3 positions were eliminated earlier from the stage of analyzing the 401-genome primary dataset. The fourth position, PB2 702, although in the first-round list, marginally failed in the subsequent validation with human entropy –0.057 and avian entropy –0.404. We proposed a computational approach capable of indicating species-associated signatures in studying human versus avian influenza viral genomes. Although we intended to analyze a comprehensive set of avian versus human influenza A viral genomes, the available sequences are predominated by H5N1 in avian viruses and H3N2 in human viruses. The short supply of sequences other than those 2 subtypes may inevitably cause a certain amount of bias in our results. At the completion of this study, we noticed a recent article by Obenauer et al., who had made 169 newly sequenced avian influenza viral genomes available to GenBank on January 26, 2006 (); these were not included in our analysis. We checked on our 52 signature positions against these new genomes and found only 2 of them that showed an entropy value slightly over our threshold –0.4. These are PB1-F2 87 and HA 237, with entropy values of –0.522, and –0.692, respectively. The choice of entropy threshold would also affect the number of signatures found. Originally we chose –0.4 on the basis of the value –0.379, computed from PB2 627 by using 95 avian genomes. We noticed that this entropy value reduced to –0.299 at PB2 627 (see Table A4) at the later validation stage, when we found 197 E and 19 K from a total of 215 avian PB2 sequences. If we chose to use a more stringent entropy threshold of –0.3, our analysis still showed 46 of those 52 reported signatures; missing were positions 73, 79, and 82 from PB1-F2, 409 from PA, and 237 and 389 from HA. In addition to the data limitations, this approach of looking for species-associated signatures by entropy is less useful for HA and NA genes. The genetic diversity that exists in either human or avian viruses for these 2 gene segments can markedly boost their respective entropy to more negative values, thus making it difficult to find residues conserved enough for identifying such signatures. We additionally performed the analysis on human H1, H2, and H3 versus avian HA (Figure A1). For NA we performed the analysis on human N1 and N2 versus avian NA. We compared 10 human H1, 3 human H2, and 293 human H3 with 95 avian HA sequences and found 13, 13, and 69 signatures (with entropy values for both human and avian within –0.4), respectively. This finding indicates that the human H1 and H2 strains are less distinct from avian strains (H5 dominant) than H3. For NA we found only 6 signatures, in comparison with 8 human N1 versus 95 avian (N1-dominant), and we found only 5 signatures when we compared 298 human N2 and 95 avian sequences. Entropy plots for these analyses can be seen in Figure A1.
Figure A1

Entropy plot for all 11 influenza proteins for human (top) versus avian (bottom). In each aligned position, we have a consensus residue for 95 avian strains displayed on top, and a consensus residue for 306 human strains at the bottom. Completely conserved amino acid positions are filled with white, while less conserved amino acids are filled in various gray shadings. Positions where one single residue dominates over 90%, less than 90% but greater than 75%, and less than 75% are labeled with red, yellow, and green letters, respectively. Yellow rectangles indicate that both human and avian flu are completely conserved to the same residue, while rectangles in magenta indicate that avian and human flu each completely conserves to a different residue Additional plots for HA, NA, NS1 and NS2, for using different counts of human or avian strains are detailed as individual captions to these plots. Adobe Acrobat PDF available at http://wwwnc.cdc.gov/eid/pdfs/06-0276-FA1.pdf (21 pages).

Two genetic alleles (allele A and B) have been described for the NS gene in avian influenza A virus. We decomposed those 95 avian NS genes into 43 in allele A and 52 in allele B and compared their amino acid sequences with 306 human NS genes. For NS1, 6 signatures were found between human viruses and avian allele A viruses, and 35 signatures were found between human viruses and avian allele B viruses. For NS2, 3 signatures were found between human viruses and allele A viruses, and 6 signatures were found between human viruses and allele B viruses. These results suggest that avian allele B viruses are more distinct from human viruses than are allele A viruses. Entropy plots and histograms for these analyses can be seen in Figure A1 and Figure A3.
Figure A3

Histograms compare 43 avian allele A viruses and 306 human viruses (panels A and C), and 52 avian allele B viruses and 306 human viruses (panels B and D), based on their NS1 and NS2 genomic segments. Vertical axis shows the count for pairs of sequences with specific percent identity (rounded to integer). Red bars represent frequencies for 'homo' pairs – sequences of the same host species (human to human, or avian to avian); blue bars represent frequencies for 'hetero' pairs – pairs that cross host species (human to avian, or avian to human). Adobe Acrobat PDF available at http://wwwnc.cdc.gov/eid/pdfs/06-0276-FA3.pdf (5 pages).

From the histograms, we found that some of the 11 genes vary greatly between human and avian viruses, while some others vary little. No boundaries were found between homopairs and heteropairs for HA, NA, and PB1 for human versus avian viruses. This finding seems reasonable because the 2 recent pandemic strains, the 1957 H2N2 and the 1968 H3N2, both originated from reassortment with avian influenza viruses (HA, NA, and PB1 gene segments were from avian influenza). On the other hand, because histograms of NP, followed by PA and PB2, may be used to distinguish human influenza viruses from avian influenza viruses, perhaps some biologic constraints against the occurrence of reassortment exist for these 3 genes. Both the M and NS genes are less differentiable between these 2 types of influenza A viruses. NP not only displays a clear boundary between human and avian viruses from histogram analysis but also contains more species-associated amino acid signatures (15 of 52) than other ORFs. In addition to NP, polymerase proteins PB2, PB1, and PA also contain abundant species-associated signatures. Most signatures in these viral RNPs are located on the functional domains related to RNP-RNP interactions that are necessary to form replicase/transcriptase complex (3P and NP), which suggests that specific combinations of polymerase complex and NP would allow an influenza virus to replicate itself efficiently (Table 1). In addition to RNA-interacting domains, many species-associated amino acid signatures of 3P and NP are located in regions related to nuclear localization signals. Influenza viral replication is highly dependent on nuclear function (), making it worthwhile to further examine the roles of those amino acid signatures on nuclear localization of viral RNP in avian versus human cells. We also noticed that several amino acid signatures in NP are located in the regions that interact with cellular proteins, such as splicing factor (BAT1/UAP56) or MxA, which plays a certain role in cellular antiviral mechanisms. What species-specific host factors may affect influenza viral replication rates is not clear. Biologic experiments are required for further understanding the roles of those amino acid residues and related functional domains in the mechanism of interspecies infection. PB1-F2 is a novel influenza viral protein translated from alternative initiation of PB1 gene. PB1-F2 of PR8 (H1N1) has been shown to target mitochondria and then trigger host cell apoptosis (). Our previous research has found that several strains contain truncated PB1-F2 (). In this study, 379 of 401 PB1 sequences (in the primary dataset) contained PB1-F2 >87 and <90 aa. For the other 22 sequences, 2 H3N2 strains missed a start codon, 3 H3N2 had the translation stopped at 11 aa, 1 H9N2 stopped at 8 aa, 5 H1N1 stopped at 57 aa, and 3 H9N2 and 7 H3N2 stopped at 79 aa. One H5N1 contained extra residues; its PB1-F2 was 101 aa. We also noted 5 species-associated signatures on PB1-F2; all of them are within the C-terminal domain, which is important for mitochondria targeting (,). Further investigation of the mitochondria localization of those PB1-F2 variants and their abilities for triggering apoptosis in cells derived from different species is warranted. How many mutations would make an avian virus capable of infecting humans efficiently, or how many mutations would render an influenza virus a pandemic strain, is difficult to predict. We have examined sequences from the 1918 strain, which is the only pandemic influenza virus that could be entirely derived from avian strains. Of the 52 species-associated positions, 16 have residues typical for human strains; the others remained as avian signatures. The result supports the hypothesis that the 1918 pandemic virus is more closely related to the avian influenza A virus than are other human influenza viruses (). From the 21 avian viruses isolated from humans in this study, we found 19 (90.5%) that contain >1 change at the species-associated sites. Upon examining signature changes from similarly sized sets of randomly selected human viruses, randomly selected avian viruses, and randomly selected viruses (avian plus human), we found 29.4%, 71.4%, and 47.1%, respectively, contain species-associated mutations. Although predicting the emergence of a pandemic strain is difficult, close monitoring of how those species-associated signature positions have changed from bird-specific to human-specific signatures may provide a measurement for the prediction of such events.

Appendix

Supporting Materials and Methods

In the main text we have mentioned an entropy value was defined at an aligned amino acid position according to the formula ΣPi*log(Pi), where i is the observed probability for each of the 20 amino acids. An entropy value defined like this is at most zero when all amino acids at this position conserve to the same residue, while a more negative value indicates that the residues are more divergent for containing more residue types. Although BioEdit also includes a module with similar formula in computing entropy values for aligned sequences, we chose to develop our own software for more streamlined data manipulation and subsequent analysis and interpretation. To reveal the host-associated amino acid signatures, we have retrieved full genome sequences (as of August 22, 2005) from the genome browser at Influenza Sequence Database. Strains containing all eight RNA segments and for each segment a minimum 90% long of the coding sequence based on PR8 were included, which serve as the primary dataset for full genome scanning. Altogether, we have 95 avian influenza genomes (including 60 H5N1, 8 H6N1, 6 H6N2, 1 H7N1, 1 H7N3, 2 H7N7, 17 H9N2) and 306 human influenza genomes (8 H1N1, 2 H1N2, 3 H2N2 and 293 H3N2), the latter include 11 complete genomes of Taiwanese strains from 1996 to 2004 (newly sequenced data from this study). See Supporting Table 1 for a complete listing of accessions for these 401 genomes. Coding sequence alignments for each genomic segment were compiled: PB2, 759 aa; PB1, 757 aa; PB1-F2, 90 aa; PA, 716 aa; HA, 591 aa; NP, 498 aa; NA, 480 aa; M1, 252 aa; M2, 97 aa; NS1, 230 aa; and NS2, 121 aa. Human-isolated avian influenza viruses from human flu were separately retrieved from NCBI as well as from ISD. Altogether we have 417 accessions from 60 avian flu strains (48 H5N1, 6 H9N2, 5 H7N7 and 1 H7N2), in which 21 strains (17 H5N1, 3 H9N2 and 1 H7N7) contain sequences (full or nearly full-length) from all 8 genomic RNAs. See Table A2 for a complete listing of these accessions. For validating the obtained signatures from analyzing the mentioned 401-genome primary dataset, we have firstly retrieved 14,057 human or avian influenza A protein sequences from NCBI's Influenza Virus Resources (as of January 17, 2006), including 5,468 avian and 8,589 human sequences (786 H1N1 sequences and 7,097 H3N2 sequences among the others). At the stage of revising this manuscript, we have included more H1N1sequences (2,514 in total, as of April 20, 2006) for validation to relieve the limitation that may be caused by the unbalanced sequence counts between H1N1 (786 sequences) and H3N2 (7,097 sequences) previously used, thus making the results more robust. Altogether we have used 15,785 influenza protein sequences for confirmatory analysis.
  35 in total

1.  Distinct regions of influenza virus PB1 polymerase subunit recognize vRNA and cRNA templates.

Authors:  S González; J Ortín
Journal:  EMBO J       Date:  1999-07-01       Impact factor: 11.598

2.  Nuclear MxA proteins form a complex with influenza virus NP and inhibit the transcription of the engineered influenza virus genome.

Authors:  Kadir Turan; Masaki Mibayashi; Kenji Sugiyama; Shoko Saito; Akiko Numajiri; Kyosuke Nagata
Journal:  Nucleic Acids Res       Date:  2004-01-29       Impact factor: 16.971

3.  Complex structure of the nuclear translocation signal of influenza virus polymerase PA subunit.

Authors:  A Nieto; S de la Luna; J Bárcena; A Portela; J Ortín
Journal:  J Gen Virol       Date:  1994-01       Impact factor: 3.891

4.  On the origin of the human influenza virus subtypes H2N2 and H3N2.

Authors:  C Scholtissek; W Rohde; V Von Hoyningen; R Rott
Journal:  Virology       Date:  1978-06-01       Impact factor: 3.616

5.  Mitochondrial targeting sequence of the influenza A virus PB1-F2 protein and its function in mitochondria.

Authors:  Hiroshi Yamada; Ritsu Chounan; Youichirou Higashi; Naoki Kurihara; Hiroshi Kido
Journal:  FEBS Lett       Date:  2004-12-17       Impact factor: 4.124

6.  The influenza virus ion channel and maturation cofactor M2 is a cholesterol-binding protein.

Authors:  Cornelia Schroeder; Harald Heider; Elisabeth Möncke-Buchner; Tse-I Lin
Journal:  Eur Biophys J       Date:  2004-06-25       Impact factor: 1.733

7.  A single amino acid in the PB2 gene of influenza A virus is a determinant of host range.

Authors:  E K Subbarao; W London; B R Murphy
Journal:  J Virol       Date:  1993-04       Impact factor: 5.103

8.  A classical bipartite nuclear localization signal on Thogoto and influenza A virus nucleoproteins.

Authors:  F Weber; G Kochs; S Gruber; O Haller
Journal:  Virology       Date:  1998-10-10       Impact factor: 3.616

9.  Identification of an RNA binding region within the N-terminal third of the influenza A virus nucleoprotein.

Authors:  C Albo; A Valencia; A Portela
Journal:  J Virol       Date:  1995-06       Impact factor: 5.103

Review 10.  Evidence of an absence: the genetic origins of the 1918 pandemic influenza virus.

Authors:  Ann H Reid; Jeffery K Taubenberger; Thomas G Fanning
Journal:  Nat Rev Microbiol       Date:  2004-11       Impact factor: 60.633

View more
  135 in total

Review 1.  Influenza A virus polymerase: structural insights into replication and host adaptation mechanisms.

Authors:  Stéphane Boivin; Stephen Cusack; Rob W H Ruigrok; Darren J Hart
Journal:  J Biol Chem       Date:  2010-06-10       Impact factor: 5.157

2.  Emergence of amantadine-resistant avian influenza H5N1 virus in India.

Authors:  C Tosh; H V Murugkar; S Nagarajan; S Tripathi; M Katare; R Jain; R Khandia; Z Syed; P Behera; S Patil; D D Kulkarni; S C Dubey
Journal:  Virus Genes       Date:  2010-10-16       Impact factor: 2.332

3.  Pathogenicity of swine influenza viruses possessing an avian or swine-origin PB2 polymerase gene evaluated in mouse and pig models.

Authors:  Wenjun Ma; Kelly M Lager; Xi Li; Bruce H Janke; Derek A Mosier; Laura E Painter; Eva S Ulery; Jingqun Ma; Porntippa Lekcharoensuk; Richard J Webby; Jürgen A Richt
Journal:  Virology       Date:  2010-11-11       Impact factor: 3.616

4.  Genetic characterization of avian influenza viruses isolated in Israel during 2000-2006.

Authors:  Natalia Golender; Alexander Panshin; Caroline Banet-Noach; Sagit Nagar; Shimon Pokamunski; Michael Pirak; Yevgeny Tendler; Irit Davidson; Maricarmen García; Shimon Perk
Journal:  Virus Genes       Date:  2008-08-20       Impact factor: 2.332

5.  An inhibitory activity in human cells restricts the function of an avian-like influenza virus polymerase.

Authors:  Andrew Mehle; Jennifer A Doudna
Journal:  Cell Host Microbe       Date:  2008-08-14       Impact factor: 21.023

6.  Evolution of an avian H5N1 influenza A virus escape mutant.

Authors:  Kamel M A Hassanin; Ahmed S Abdel-Moneim
Journal:  World J Virol       Date:  2013-11-12

Review 7.  Pathogenicity of highly pathogenic avian influenza virus in mammals.

Authors:  Emmie de Wit; Yoshihiro Kawaoka; Menno D de Jong; Ron A M Fouchier
Journal:  Vaccine       Date:  2008-09-12       Impact factor: 3.641

8.  Human and avian influenza viruses target different cells in the lower respiratory tract of humans and other mammals.

Authors:  Debby van Riel; Vincent J Munster; Emmie de Wit; Guus F Rimmelzwaan; Ron A M Fouchier; Albert D M E Osterhaus; Thijs Kuiken
Journal:  Am J Pathol       Date:  2007-08-23       Impact factor: 4.307

9.  Identification of amino acid changes that may have been critical for the genesis of A(H7N9) influenza viruses.

Authors:  Gabriele Neumann; Catherine A Macken; Yoshihiro Kawaoka
Journal:  J Virol       Date:  2014-02-12       Impact factor: 5.103

10.  Role of host-specific amino acids in the pathogenicity of avian H5N1 influenza viruses in mice.

Authors:  Jin Hyun Kim; Masato Hatta; Shinji Watanabe; Gabriele Neumann; Tokiko Watanabe; Yoshihiro Kawaoka
Journal:  J Gen Virol       Date:  2009-12-16       Impact factor: 3.891

View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.