| Literature DB >> 32133509 |
Kwang Su Jung1, Kyung-Won Hong1, Hyun Youn Jo1, Jongpill Choi1, Hyo-Jeong Ban2, Seong Beom Cho1, Myungguen Chung1.
Abstract
Since 2012, the Center for Genome Science of the Korea National Institute of Health (KNIH) has been sequencing complete genomes of 1722 Korean individuals. As a result, more than 32 million variant sites have been identified, and a large proportion of the variant sites have been detected for the first time. In this article, we describe the Korean Reference Genome Database (KRGDB) and its genome browser. The current version of our database contains both single nucleotide and short insertion/deletion variants. The DNA samples were obtained from four different origins and sequenced in different sequencing depths (10× coverage of 63 individuals, 20× coverage of 194 individuals, combined 10× and 20× coverage of 135 individuals, 30× coverage of 230 individuals and 30× coverage of 1100 individuals). The major features of the KRGDB are that it contains information on the Korean genomic variant frequency, frequency difference between the Korean and other populations and the variant functional annotation (such as regulatory elements in ENCODE regions and coding variant functions) of the variant sites. Additionally, we performed the genome-wide association study (GWAS) between Korean genome variant sites for the 30×230 individuals and three major common diseases (diabetes, hypertension and metabolic syndrome). The association results are displayed on our browser. The KRGDB uses the MySQL database and Apache-Tomcat web server adopted with Java Server Page (JSP) and is freely available at http://coda.nih.go.kr/coda/KRGDB/index.jsp. Availability: http://coda.nih.go.kr/coda/KRGDB/index.jsp.Entities:
Mesh:
Year: 2020 PMID: 32133509 PMCID: PMC7056612 DOI: 10.1093/database/baz146
Source DB: PubMed Journal: Database (Oxford) ISSN: 1758-0463 Impact factor: 3.451
The KRG individual groups
|
|
|
|
|
|
|---|---|---|---|---|
| The first phase (2012–2014) | 63 | Korea National Health and Nutrition Examination Survey | 10× | HiSeq 2000 |
| 194 | Volunteers who participated in the Korean Genome Organization Conference | 20× | ||
| 230 | The Ansan-Ansung cohort (epidemiological and genotype data) | 30× | ||
| 135 | The Ansan-Ansung cohort (genotype) : merged 30× (10× in 2012 and 20× in 2013) | 30× (10×+20×) | ||
| The second phase (2015–2016) | 1100 | The Korean Biobank Project | 30× | HiSeq X Ten |
Figure 1System architecture of KRGDB and Genome Browser. The system mainly consists of variation/annotation database and its genome browser.
Figure 2Alternative allele frequency difference between KRG and HapMap III ethnics. The horizontal axis denotes the genomic positions of the chosen chromosome (chr1).
Figure 3Alternative allele frequency difference between KRG and 1000 Genomes ethnics. The horizontal axis denotes the genomic positions of the chosen chromosome (chr1).
Figure 4Disease risks of type II diabetes (DM), hypertension (HTN) and metabolic syndrome (MS). Each dot represents risk P values (–logP). The red and blue colour indicates odds ratio ≥1.0 and <1.0, respectively. The horizontal axis denotes the genomic positions of the chosen chromosome (chr1).