| Literature DB >> 33920487 |
Miao Miao1, Erik De Clercq2, Guangdi Li1.
Abstract
Severe acute respiratory syndrome coronavirus 2 (Entities:
Keywords: COVID-19; SARS-CoV-2; genetic diversity; genetic variant; global pandemic
Year: 2021 PMID: 33920487 PMCID: PMC8069977 DOI: 10.3390/biomedicines9040412
Source DB: PubMed Journal: Biomedicines ISSN: 2227-9059
Figure 1Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) genome organization and the distribution of the SARS-CoV-2 genome sequence used in this study. (A) Genomic architecture of SARS-CoV-2 based on the reference sequence (Wuhan-Hu-1, NCBI accession NC_045512). (B) The temporal and geographic distribution of all sequences. All temporal analyses were based on the date of sequence collection.
Basic characteristics of SARS-CoV-2 proteins.
| Protein | Gene | Protein Location | Protein Length | Polymorphic Sites | Substitution Rate (%) |
|---|---|---|---|---|---|
| NSP1 |
| 266–805 | 180 | 177 | 0.03 |
| NSP2 |
| 806–2719 | 638 | 608 | 0.07 |
| NSP3 |
| 2720–8554 | 1945 | 1752 | 0.03 |
| NSP4 |
| 8555–10,054 | 500 | 419 | 0.02 |
| 3C-like protease (NSP5) |
| 10,055–10,972 | 306 | 251 | 0.04 |
| NSP6 |
| 10,973–11,842 | 290 | 256 | 0.06 |
| NSP7 |
| 11,843–12,091 | 83 | 72 | 0.04 |
| NSP8 |
| 12,092–12,685 | 198 | 176 | 0.03 |
| NSP9 |
| 12,686–13,024 | 113 | 93 | 0.03 |
| NSP10 |
| 13,025–13,441 | 139 | 105 | 0.01 |
| NSP11 |
| 13,442–13,480 | 13 | 11 | 0.01 |
| RNA-dependent RNA polymerase (NSP12) |
| 13,442–13,468 13,468–16,236 | 932 | 715 | 0.13 |
| Helicase (NSP13) |
| 16,237–18,039 | 601 | 478 | 0.04 |
| 3′-to-5′ exonuclease (NSP14) |
| 18,040–19,620 | 527 | 444 | 0.03 |
| endoRNAse (NSP15) |
| 19,621–20,658 | 346 | 313 | 0.04 |
| 2′-O-ribose methyltransferase (NSP16) |
| 20,659–21,552 | 298 | 245 | 0.03 |
| Spike glycoprotein (S) |
| 21,563–25,384 | 1273 | 1096 | 0.14 |
| ORF3a |
| 25,393–26,220 | 275 | 273 | 0.20 |
| Envelope protein (E) |
| 26,245–26,472 | 75 | 72 | 0.02 |
| Membrane protein (M) |
| 26,523–27,191 | 222 | 168 | 0.02 |
| ORF6 |
| 27,202–27,387 | 61 | 60 | 0.03 |
| ORF7a |
| 27,394–27,759 | 121 | 121 | 0.04 |
| ORF7b |
| 27,756–27,887 | 43 | 43 | 0.07 |
| ORF8 |
| 27,894–28,259 | 121 | 121 | 0.12 |
| Nucleocapsid protein (N) |
| 28,274–29,533 | 419 | 378 | 0.31 |
| ORF10 |
| 29,558–29,674 | 38 | 37 | 0.72 |
Figure 2Temporal distributions of the substitution rates of the six SARS-CoV-2 proteins. The proteins were the top six proteins with the highest substitution rates. (A) The global substitution rate curves of SARS-CoV-2 proteins based on a time sliding window. (B) The substitution rate curves of SARS-CoV-2 proteins based on a time sliding window on different continents. The vertical axis represents the moving-window substitution rate, calculated by dividing the total number of polymorphic sites of a protein contained in the sequences, 15 days before and after a specific date, by the total number of all positions of the protein in the period.
Figure 3Distribution of variants at positions of SARS-CoV-2 proteins. For each site, the reference index is shown at the top, followed by variants with a frequency >1%. Variants highlighted with green superscripts were those with frequencies >5%.
Figure 4Distribution of frequent variants across the SARS-CoV-2 proteins in different geographic areas. (A) Frequencies of the top 30 variants, with the highest variant frequencies in different continents. (B) Temporal and geographic dynamics of frequent SARS-CoV-2 variants. The vertical axis represents the moving-window variant frequency calculated by dividing the number of sequences containing a specific variant, 15 days before and after a specific date, by the total number of sequences in the period (31 days).
The overall prevalence of SARS-CoV-2 variants in a large-scale dataset of 260,673 sequences.
| Genetic Variant | Variant Frequency (%) | ||||||
|---|---|---|---|---|---|---|---|
| Asia | Africa | Europe | North America | South America | Oceania | Total | |
| T85I(NSP2) | 10.23 | 5.72 | 3.26 | 52.63 | 8.84 | 5.96 | 15.38 |
| I120F(NSP2) | 3.43 | 0.00 | 2.12 | 0.02 | 0.00 | 72.98 | 5.23 |
| M324I(NSP4) | 0.25 | 0.41 | 4.38 | 0.08 | 0.09 | 0.11 | 2.84 |
| L89F(NSP5) | 0.23 | 0.03 | 0.08 | 15.73 | 0.06 | 0.33 | 3.75 |
| L37F(NSP6) | 19.93 | 4.93 | 6.84 | 3.28 | 3.09 | 4.78 | 6.52 |
| A185S(NSP12) | 0.15 | 0.34 | 4.35 | 0.04 | 0.00 | 0.12 | 2.81 |
| P323L(NSP12) | 71.54 | 91.83 | 96.04 | 92.87 | 97.29 | 91.22 | 93.74 |
| V776L(NSP12) | 0.18 | 0.34 | 4.33 | 0.24 | 0.00 | 0.12 | 2.84 |
| K218R(NSP13) | 0.15 | 0.34 | 4.32 | 0.02 | 0.00 | 0.11 | 2.78 |
| E261D(NSP13) | 0.20 | 0.58 | 4.39 | 0.08 | 0.03 | 0.11 | 2.84 |
| H290Y(NSP13) | 0.55 | 0.17 | 3.80 | 0.25 | 0.03 | 0.22 | 2.53 |
| N129D(NSP14) | 0.17 | 0.00 | 0.05 | 12.61 | 0.03 | 0.32 | 3.00 |
| R216C(NSP16) | 0.20 | 0.00 | 0.07 | 12.40 | 0.03 | 0.33 | 2.96 |
| L18F(S) | 0.32 | 0.28 | 18.78 | 0.30 | 0.34 | 0.12 | 12.08 |
| A222V(S) | 0.50 | 0.79 | 40.78 | 0.17 | 0.34 | 0.53 | 26.14 |
| N439K(S) | 0.45 | 0.03 | 3.90 | 0.01 | 0.06 | 0.06 | 2.52 |
| S477N(S) | 0.19 | 0.66 | 4.82 | 0.14 | 0.12 | 71.22 | 6.62 |
| D614G(S) | 71.91 | 93.91 | 96.11 | 93.10 | 97.26 | 91.26 | 93.88 |
| Q57H(ORF3a) | 24.97 | 12.34 | 11.86 | 59.86 | 14.94 | 8.08 | 23.59 |
| G172V(ORF3a) | 0.20 | 0.07 | 0.07 | 12.07 | 0.06 | 0.31 | 2.89 |
| S24L(ORF8) | 0.23 | 0.17 | 0.15 | 21.78 | 0.03 | 0.90 | 5.24 |
| P67S(N) | 0.17 | 0.03 | 0.11 | 12.21 | 0.06 | 0.33 | 2.94 |
| S194L(N) | 9.64 | 6.54 | 4.27 | 12.50 | 0.19 | 1.19 | 6.29 |
| P199L(N) | 0.29 | 0.21 | 2.29 | 12.49 | 0.06 | 0.40 | 4.42 |
| R203K(N) | 36.96 | 52.15 | 28.16 | 13.35 | 69.12 | 78.12 | 28.45 |
| G204R(N) | 36.73 | 51.62 | 27.71 | 13.31 | 69.07 | 78.08 | 28.13 |
| A220V(N) | 0.45 | 1.10 | 40.53 | 0.13 | 0.09 | 0.35 | 25.96 |
| M234I(N) | 0.50 | 0.55 | 4.77 | 0.61 | 3.64 | 0.30 | 3.28 |
| A376T(N) | 0.15 | 0.35 | 4.33 | 0.02 | 0.00 | 0.11 | 2.78 |
| V30L(ORF10) | 0.58 | 0.69 | 40.64 | 0.08 | 0.13 | 0.35 | 26.01 |
Figure 5Variants in the SARS-CoV-2 proteins. (A) Site 222 (magenta) and Site 614 (red) of spike. Almost all sequences showed a variant (D614G). Site 614 is located at the interface between two subunits. (B) Site 323 (deep salmon) of the NSP12 protein. Many sequences showed a variant (P323L). (C) Site 203 (orange), Site 204 (hot pink), and Site 220 (lime green) of the nucleocapsid protein. The frequencies of variants R203K, G204R, and A220V were high. (D) Site 85 (light pink) of the NSP2 protein. Many sequences showed a variant (T85I). The structures of spike, NSP12, the nucleocapsid protein, and NSP2 were collected from https://zhanglab.ccmb.med.umich.edu/COVID-19/. The protein structural figures were generated by the software PyMOL (http://www.pymol.org/, the accessed date: 16 January 2021).
Figure 6The substitution matrix at the nucleic acid level. Nucleotide substitutions showed in this figure are the main substitutions characterizing the SARS-CoV-2 clades.
Figure 7Temporal distributions of the 12 clades based on 260,673 complete SARS-CoV-2 nucleotide sequences. (A) The global distribution of the 12 clades over time. (B) The distribution of the 12 clades over time on six continents. The vertical axis represents the moving-window proportion, calculated by dividing the number of sequences belonging to a specific clade, 15 days before and after a specific date, by the total number of sequences in the period (31 days).