Literature DB >> 19958513

Gevab: a prototype genome variation analysis browsing server.

Woo-Yeon Kim1, Sang-Yoon Kim, Tae-Hyung Kim, Sung-Min Ahn, Ha Na Byun, Deokhoon Kim, Dae-Soo Kim, Yong Seok Lee, Ho Ghang, Daeui Park, Byoung-Chul Kim, Chulhong Kim, Sunghoon Lee, Seong-Jin Kim, Jong Bhak.   

Abstract

BACKGROUND: The first Korean individual diploid genome sequence data (KOREF) was publicized in December 2008.
RESULTS: A Korean genome variation analysis and browsing server (Gevab) was constructed as a database and web server for the exploration and downloading of Korean personal genome(s). Information in the Gevab includes SNPs, short indels, and structural variation (SV) and comparison analysis between the NCBI human reference and the Korean genome(s). The user can find information on assembled consensus sequences, sequenced short reads, genetic variations, and relationships between genotype and phenotypes.
CONCLUSION: This server is openly and publicly available online at http://koreagenome.org/en/ or directly http://gevab.org.

Entities:  

Mesh:

Year:  2009        PMID: 19958513      PMCID: PMC2788354          DOI: 10.1186/1471-2105-10-S15-S3

Source DB:  PubMed          Journal:  BMC Bioinformatics        ISSN: 1471-2105            Impact factor:   3.169


Background

Most known genome browsers, such as NCBI genome [1] and Craig Venter's genome browsers [2], were built for consensus sequences from multiple individuals to construct a reference human genome. Examples of haplotype genome browsers are NCBI, UCSC [3], Ensembl [4], and Venter genome browsers. Recently, the first Asian (Chinese) diploid genome database was published, containing analysis and browsing facilities [5,6]. There are a number of general purpose genome annotation servers. They include Entrez Gene [7], Ensembl genes, OMIM [8] disease associations, HapMap [9], SNPedia [10], and genetic variations of several individual genomes such as Venter [11], Watson [12], YH (Chinese), and NA18507 (Yoruba) [13]. We have developed an individual genome variation analysis and browsing server (Gevab) for the first Korean personal genome sequence (KOREF). This server is useful to analyze a diploid human genome produced to study the complex features of human genetic variations. The system integrated multiple variation information such as Venter, Watson, YH, dbSNP, and HapMap genotypes as well as gene information. Hence, users can comparatively study the genotypes in human. Gevab also provides information for SNPs, short indels, and SVs on the KOREF genome. Gevab has two parts: genome variation analysis and genome mapping.

Materials and methods

Data source

KOREF data were generated using the Illumina GA and resulted in 82.73 gigabase (Gb) of sequence (about 1248 million paired 36-base reads and about 504 million 75-base reads). Using the MAQ (Mapping and Assembly with Qualities) [14] program, these sequences were aligned to the NCBI human genome reference (build 36, without Ns, 2,858,029,377 bp). In total, 99.9% of the NCBI reference genome was covered with an average of 25.92-fold depth (sequencing depth was 28.95-fold).

Database and browser software

In the Gevab Korean genome variation browsing part, the consensus genome sequence and genetic variants include SNPs, short indels, and SVs can be displayed. Gevab used GBrowse [15] developed by GMOD [16] for variation viewing, and the genome map browser part was developed by KOBIC.

Analysis of KOREF

From the KOREF genome sequence, 3.44 millions SNPs were identified and validated using Illumina 1 M-duo and Affy 6.0 BeadChip. We identified 342,965 short indels (-29 - +14 bp). Indels that co-occurred within a window size of 20 bp were filtered out, since they were primarily from length polymorphisms in homopolymeric tracts of A or T. Using paired-end reads, we found 2920 deletions and 415 inversion structural variants (SV) in the range of 0.1~100 kb. In addition, we detected 963 insertion events in the range of 175~250 bp. These insertions are present in the KOREF genome but absent in the NCBI reference genome. MySql and PHP, python, and AJAX were used in database construction and interface utility.

Results

Features of Gevab

The Gevab has genome variation analysis and genome map browser parts. The genome variation analysis part contains external public data sources, including the reference sequence of the human genome ((NCBI build 36), the Ensembl gene annotation, the Entrez gene annotations, dbSNP ver. 129 [17], OMIM annotations, and SNP frequencies of the HapMap population as well as genotype, indel, and structure variation of the KOREF. It is also integrated with other individual SNP variants such as James Watson's, Craig Venter's, and YangHuang's genotypes (Table 1). These external data sets are coordinated with the NCBI reference genome. A search can be done by putting in a genome location, a gene symbol, a RefSeq id, a dbSNP id, or an Ensembl gene id. When the user searches Gevab with a query, a graphical view of a chromosome ideogram and contigs are displayed. The gene locations within the 2 MB region centered on the query are also represented. For the displayed region in our browser, users can also download data with gff or fasta format ftp://ftp.kobic.re.kr/pub/KOBIC-KoreanGenome/.
Table 1

Features of Gevab, Venter, Watson, and YH genome browsers. Availability of features is indicated by "O" for "yes" and "X" for "no."

GevabVenterWatsonYH
genomeKoreanCaucasianCaucasianChinese
read mappingOXXO
sequencing coverageOOOX
genotypeOOOO
indelOOXX
structure variationOOXX

variations to compareVenter, Watson, YH, dbSNP, HapMapdbSNPdbSNP, HapMapdbSNP, HapMap

web sitehttp://gevab.orghttp://huref.jcvi.orghttp://jimwatsonsequence.cshl.eduhttp://yh.genomics.org.cn/
Features of Gevab, Venter, Watson, and YH genome browsers. Availability of features is indicated by "O" for "yes" and "X" for "no."

Gevab's map browser

The genome map browser provides reads mapping and quality information obtained from a personal genome project. A search can be done by chromosomal position. The width of a displayed region can be controlled. The browser also has zoom in and out and left and right movement functions. When a user a chooses 1000 bp window size or longer, the browser shows a graphical view with forward and reverse pair-end reads in dark and light green, and single reads in red (Figure 1(A)). For shorter than 1000 bp window size, the browser is converted to a text mode that additionally shows quality information of mapped reads (Figure 1(B)).
Figure 1

A screenshot of the genome map browser. (A) Graphic mode (B) Text mode.

A screenshot of the genome map browser. (A) Graphic mode (B) Text mode.

Gevab's variation browser

To study genome variations, a genome variation browser is more useful than a genome map browser. As an example, if a user is interested in the "NOC2L" gene, s/he can get KOREF, Watson, YH, and Venter genome variation information through the variation browser part (Figure 2).
Figure 2

A screenshot of the genome variation browser.

A screenshot of the genome variation browser.

KOREF data access

The KOREF database is developed and maintained by KOBIC (Korean Bioinformation Center). The database contains all the raw and processed data of KOREF, including KOREF consensus sequence, genetic variants, and short read alignments. These data are available for downloading. The KOREF data have been deposited in the NCBI Short Read Archive (Accession Number SRA008175).

Conclusion

Gevab contains all the raw and processed data of a Korean genome sequence, variants, and annotation. Gevab provides open and public access to all data of an individual personal diploid genome. The variation browser part was designed to present genetic variant evidence, including the position, number, and status of reads, GC content, and several mapping information. These provide valuable detailed information such as comparison and validation of genetic variations to further communities for sequencing individual genome.

Competing interests

The authors declare that they have no competing interests.

Authors' contributions

WYK developed the genetic variation browser and helped to write manuscript. SYK developed the genome map browser. THK wrote manuscript. SMA and DK provided the KOREF data. HNB, DSK, YSL, HG, DP, BCK, CK, and SL provided counseling on issues related to Gevab development. SJK supervised the whole project and guided to production of KOREF data. JB supervised the bioinformatic analysis and manuscript writing.

Note

Other papers from the meeting have been published as part of BMC Genomics Volume 10 Supplement 3, 2009: Eighth International Conference on Bioinformatics (InCoB2009): Computational Biology, available online at http://www.biomedcentral.com/1471-2164/10?issue=S3.
  13 in total

1.  dbSNP: the NCBI database of genetic variation.

Authors:  S T Sherry; M H Ward; M Kholodov; J Baker; L Phan; E M Smigielski; K Sirotkin
Journal:  Nucleic Acids Res       Date:  2001-01-01       Impact factor: 16.971

2.  The generic genome browser: a building block for a model organism system database.

Authors:  Lincoln D Stein; Christopher Mungall; ShengQiang Shu; Michael Caudy; Marco Mangone; Allen Day; Elizabeth Nickerson; Jason E Stajich; Todd W Harris; Adrian Arva; Suzanna Lewis
Journal:  Genome Res       Date:  2002-10       Impact factor: 9.043

3.  The complete genome of an individual by massively parallel DNA sequencing.

Authors:  David A Wheeler; Maithreyan Srinivasan; Michael Egholm; Yufeng Shen; Lei Chen; Amy McGuire; Wen He; Yi-Ju Chen; Vinod Makhijani; G Thomas Roth; Xavier Gomes; Karrie Tartaro; Faheem Niazi; Cynthia L Turcotte; Gerard P Irzyk; James R Lupski; Craig Chinault; Xing-zhi Song; Yue Liu; Ye Yuan; Lynne Nazareth; Xiang Qin; Donna M Muzny; Marcel Margulies; George M Weinstock; Richard A Gibbs; Jonathan M Rothberg
Journal:  Nature       Date:  2008-04-17       Impact factor: 49.962

4.  The diploid genome sequence of an Asian individual.

Authors:  Jun Wang; Wei Wang; Ruiqiang Li; Yingrui Li; Geng Tian; Laurie Goodman; Wei Fan; Junqing Zhang; Jun Li; Juanbin Zhang; Yiran Guo; Binxiao Feng; Heng Li; Yao Lu; Xiaodong Fang; Huiqing Liang; Zhenglin Du; Dong Li; Yiqing Zhao; Yujie Hu; Zhenzhen Yang; Hancheng Zheng; Ines Hellmann; Michael Inouye; John Pool; Xin Yi; Jing Zhao; Jinjie Duan; Yan Zhou; Junjie Qin; Lijia Ma; Guoqing Li; Zhentao Yang; Guojie Zhang; Bin Yang; Chang Yu; Fang Liang; Wenjie Li; Shaochuan Li; Dawei Li; Peixiang Ni; Jue Ruan; Qibin Li; Hongmei Zhu; Dongyuan Liu; Zhike Lu; Ning Li; Guangwu Guo; Jianguo Zhang; Jia Ye; Lin Fang; Qin Hao; Quan Chen; Yu Liang; Yeyang Su; A San; Cuo Ping; Shuang Yang; Fang Chen; Li Li; Ke Zhou; Hongkun Zheng; Yuanyuan Ren; Ling Yang; Yang Gao; Guohua Yang; Zhuo Li; Xiaoli Feng; Karsten Kristiansen; Gane Ka-Shu Wong; Rasmus Nielsen; Richard Durbin; Lars Bolund; Xiuqing Zhang; Songgang Li; Huanming Yang; Jian Wang
Journal:  Nature       Date:  2008-11-06       Impact factor: 49.962

5.  A second generation human haplotype map of over 3.1 million SNPs.

Authors:  Kelly A Frazer; Dennis G Ballinger; David R Cox; David A Hinds; Laura L Stuve; Richard A Gibbs; John W Belmont; Andrew Boudreau; Paul Hardenbol; Suzanne M Leal; Shiran Pasternak; David A Wheeler; Thomas D Willis; Fuli Yu; Huanming Yang; Changqing Zeng; Yang Gao; Haoran Hu; Weitao Hu; Chaohua Li; Wei Lin; Siqi Liu; Hao Pan; Xiaoli Tang; Jian Wang; Wei Wang; Jun Yu; Bo Zhang; Qingrun Zhang; Hongbin Zhao; Hui Zhao; Jun Zhou; Stacey B Gabriel; Rachel Barry; Brendan Blumenstiel; Amy Camargo; Matthew Defelice; Maura Faggart; Mary Goyette; Supriya Gupta; Jamie Moore; Huy Nguyen; Robert C Onofrio; Melissa Parkin; Jessica Roy; Erich Stahl; Ellen Winchester; Liuda Ziaugra; David Altshuler; Yan Shen; Zhijian Yao; Wei Huang; Xun Chu; Yungang He; Li Jin; Yangfan Liu; Yayun Shen; Weiwei Sun; Haifeng Wang; Yi Wang; Ying Wang; Xiaoyan Xiong; Liang Xu; Mary M Y Waye; Stephen K W Tsui; Hong Xue; J Tze-Fei Wong; Luana M Galver; Jian-Bing Fan; Kevin Gunderson; Sarah S Murray; Arnold R Oliphant; Mark S Chee; Alexandre Montpetit; Fanny Chagnon; Vincent Ferretti; Martin Leboeuf; Jean-François Olivier; Michael S Phillips; Stéphanie Roumy; Clémentine Sallée; Andrei Verner; Thomas J Hudson; Pui-Yan Kwok; Dongmei Cai; Daniel C Koboldt; Raymond D Miller; Ludmila Pawlikowska; Patricia Taillon-Miller; Ming Xiao; Lap-Chee Tsui; William Mak; You Qiang Song; Paul K H Tam; Yusuke Nakamura; Takahisa Kawaguchi; Takuya Kitamoto; Takashi Morizono; Atsushi Nagashima; Yozo Ohnishi; Akihiro Sekine; Toshihiro Tanaka; Tatsuhiko Tsunoda; Panos Deloukas; Christine P Bird; Marcos Delgado; Emmanouil T Dermitzakis; Rhian Gwilliam; Sarah Hunt; Jonathan Morrison; Don Powell; Barbara E Stranger; Pamela Whittaker; David R Bentley; Mark J Daly; Paul I W de Bakker; Jeff Barrett; Yves R Chretien; Julian Maller; Steve McCarroll; Nick Patterson; Itsik Pe'er; Alkes Price; Shaun Purcell; Daniel J Richter; Pardis Sabeti; Richa Saxena; Stephen F Schaffner; Pak C Sham; Patrick Varilly; David Altshuler; Lincoln D Stein; Lalitha Krishnan; Albert Vernon Smith; Marcela K Tello-Ruiz; Gudmundur A Thorisson; Aravinda Chakravarti; Peter E Chen; David J Cutler; Carl S Kashuk; Shin Lin; Gonçalo R Abecasis; Weihua Guan; Yun Li; Heather M Munro; Zhaohui Steve Qin; Daryl J Thomas; Gilean McVean; Adam Auton; Leonardo Bottolo; Niall Cardin; Susana Eyheramendy; Colin Freeman; Jonathan Marchini; Simon Myers; Chris Spencer; Matthew Stephens; Peter Donnelly; Lon R Cardon; Geraldine Clarke; David M Evans; Andrew P Morris; Bruce S Weir; Tatsuhiko Tsunoda; James C Mullikin; Stephen T Sherry; Michael Feolo; Andrew Skol; Houcan Zhang; Changqing Zeng; Hui Zhao; Ichiro Matsuda; Yoshimitsu Fukushima; Darryl R Macer; Eiko Suda; Charles N Rotimi; Clement A Adebamowo; Ike Ajayi; Toyin Aniagwu; Patricia A Marshall; Chibuzor Nkwodimmah; Charmaine D M Royal; Mark F Leppert; Missy Dixon; Andy Peiffer; Renzong Qiu; Alastair Kent; Kazuto Kato; Norio Niikawa; Isaac F Adewole; Bartha M Knoppers; Morris W Foster; Ellen Wright Clayton; Jessica Watkin; Richard A Gibbs; John W Belmont; Donna Muzny; Lynne Nazareth; Erica Sodergren; George M Weinstock; David A Wheeler; Imtaz Yakub; Stacey B Gabriel; Robert C Onofrio; Daniel J Richter; Liuda Ziaugra; Bruce W Birren; Mark J Daly; David Altshuler; Richard K Wilson; Lucinda L Fulton; Jane Rogers; John Burton; Nigel P Carter; Christopher M Clee; Mark Griffiths; Matthew C Jones; Kirsten McLay; Robert W Plumb; Mark T Ross; Sarah K Sims; David L Willey; Zhu Chen; Hua Han; Le Kang; Martin Godbout; John C Wallenburg; Paul L'Archevêque; Guy Bellemare; Koji Saeki; Hongguang Wang; Daochang An; Hongbo Fu; Qing Li; Zhen Wang; Renwu Wang; Arthur L Holden; Lisa D Brooks; Jean E McEwen; Mark S Guyer; Vivian Ota Wang; Jane L Peterson; Michael Shi; Jack Spiegel; Lawrence M Sung; Lynn F Zacharia; Francis S Collins; Karen Kennedy; Ruth Jamieson; John Stewart
Journal:  Nature       Date:  2007-10-18       Impact factor: 49.962

6.  Entrez Gene: gene-centered information at NCBI.

Authors:  Donna Maglott; Jim Ostell; Kim D Pruitt; Tatiana Tatusova
Journal:  Nucleic Acids Res       Date:  2006-12-05       Impact factor: 16.971

7.  The diploid genome sequence of an individual human.

Authors:  Samuel Levy; Granger Sutton; Pauline C Ng; Lars Feuk; Aaron L Halpern; Brian P Walenz; Nelson Axelrod; Jiaqi Huang; Ewen F Kirkness; Gennady Denisov; Yuan Lin; Jeffrey R MacDonald; Andy Wing Chun Pang; Mary Shago; Timothy B Stockwell; Alexia Tsiamouri; Vineet Bafna; Vikas Bansal; Saul A Kravitz; Dana A Busam; Karen Y Beeson; Tina C McIntosh; Karin A Remington; Josep F Abril; John Gill; Jon Borman; Yu-Hui Rogers; Marvin E Frazier; Stephen W Scherer; Robert L Strausberg; J Craig Venter
Journal:  PLoS Biol       Date:  2007-09-04       Impact factor: 8.029

8.  McKusick's Online Mendelian Inheritance in Man (OMIM).

Authors:  Joanna Amberger; Carol A Bocchini; Alan F Scott; Ada Hamosh
Journal:  Nucleic Acids Res       Date:  2008-10-08       Impact factor: 16.971

9.  The HuRef Browser: a web resource for individual human genomics.

Authors:  Nelson Axelrod; Yuan Lin; Pauline C Ng; Timothy B Stockwell; Jonathan Crabtree; Jiaqi Huang; Ewen Kirkness; Robert L Strausberg; Marvin E Frazier; J Craig Venter; Saul Kravitz; Samuel Levy
Journal:  Nucleic Acids Res       Date:  2008-11-26       Impact factor: 16.971

10.  Accurate whole human genome sequencing using reversible terminator chemistry.

Authors:  David R Bentley; Shankar Balasubramanian; Harold P Swerdlow; Geoffrey P Smith; John Milton; Clive G Brown; Kevin P Hall; Dirk J Evers; Colin L Barnes; Helen R Bignell; Jonathan M Boutell; Jason Bryant; Richard J Carter; R Keira Cheetham; Anthony J Cox; Darren J Ellis; Michael R Flatbush; Niall A Gormley; Sean J Humphray; Leslie J Irving; Mirian S Karbelashvili; Scott M Kirk; Heng Li; Xiaohai Liu; Klaus S Maisinger; Lisa J Murray; Bojan Obradovic; Tobias Ost; Michael L Parkinson; Mark R Pratt; Isabelle M J Rasolonjatovo; Mark T Reed; Roberto Rigatti; Chiara Rodighiero; Mark T Ross; Andrea Sabot; Subramanian V Sankar; Aylwyn Scally; Gary P Schroth; Mark E Smith; Vincent P Smith; Anastassia Spiridou; Peta E Torrance; Svilen S Tzonev; Eric H Vermaas; Klaudia Walter; Xiaolin Wu; Lu Zhang; Mohammed D Alam; Carole Anastasi; Ify C Aniebo; David M D Bailey; Iain R Bancarz; Saibal Banerjee; Selena G Barbour; Primo A Baybayan; Vincent A Benoit; Kevin F Benson; Claire Bevis; Phillip J Black; Asha Boodhun; Joe S Brennan; John A Bridgham; Rob C Brown; Andrew A Brown; Dale H Buermann; Abass A Bundu; James C Burrows; Nigel P Carter; Nestor Castillo; Maria Chiara E Catenazzi; Simon Chang; R Neil Cooley; Natasha R Crake; Olubunmi O Dada; Konstantinos D Diakoumakos; Belen Dominguez-Fernandez; David J Earnshaw; Ugonna C Egbujor; David W Elmore; Sergey S Etchin; Mark R Ewan; Milan Fedurco; Louise J Fraser; Karin V Fuentes Fajardo; W Scott Furey; David George; Kimberley J Gietzen; Colin P Goddard; George S Golda; Philip A Granieri; David E Green; David L Gustafson; Nancy F Hansen; Kevin Harnish; Christian D Haudenschild; Narinder I Heyer; Matthew M Hims; Johnny T Ho; Adrian M Horgan; Katya Hoschler; Steve Hurwitz; Denis V Ivanov; Maria Q Johnson; Terena James; T A Huw Jones; Gyoung-Dong Kang; Tzvetana H Kerelska; Alan D Kersey; Irina Khrebtukova; Alex P Kindwall; Zoya Kingsbury; Paula I Kokko-Gonzales; Anil Kumar; Marc A Laurent; Cynthia T Lawley; Sarah E Lee; Xavier Lee; Arnold K Liao; Jennifer A Loch; Mitch Lok; Shujun Luo; Radhika M Mammen; John W Martin; Patrick G McCauley; Paul McNitt; Parul Mehta; Keith W Moon; Joe W Mullens; Taksina Newington; Zemin Ning; Bee Ling Ng; Sonia M Novo; Michael J O'Neill; Mark A Osborne; Andrew Osnowski; Omead Ostadan; Lambros L Paraschos; Lea Pickering; Andrew C Pike; Alger C Pike; D Chris Pinkard; Daniel P Pliskin; Joe Podhasky; Victor J Quijano; Come Raczy; Vicki H Rae; Stephen R Rawlings; Ana Chiva Rodriguez; Phyllida M Roe; John Rogers; Maria C Rogert Bacigalupo; Nikolai Romanov; Anthony Romieu; Rithy K Roth; Natalie J Rourke; Silke T Ruediger; Eli Rusman; Raquel M Sanches-Kuiper; Martin R Schenker; Josefina M Seoane; Richard J Shaw; Mitch K Shiver; Steven W Short; Ning L Sizto; Johannes P Sluis; Melanie A Smith; Jean Ernest Sohna Sohna; Eric J Spence; Kim Stevens; Neil Sutton; Lukasz Szajkowski; Carolyn L Tregidgo; Gerardo Turcatti; Stephanie Vandevondele; Yuli Verhovsky; Selene M Virk; Suzanne Wakelin; Gregory C Walcott; Jingwen Wang; Graham J Worsley; Juying Yan; Ling Yau; Mike Zuerlein; Jane Rogers; James C Mullikin; Matthew E Hurles; Nick J McCooke; John S West; Frank L Oaks; Peter L Lundberg; David Klenerman; Richard Durbin; Anthony J Smith
Journal:  Nature       Date:  2008-11-06       Impact factor: 49.962

View more
  2 in total

1.  Towards a career in bioinformatics.

Authors:  Shoba Ranganathan
Journal:  BMC Bioinformatics       Date:  2009-12-03       Impact factor: 3.169

2.  Whole genome sequencing of 35 individuals provides insights into the genetic architecture of Korean population.

Authors:  Wenqian Zhang; Joe Meehan; Zhenqiang Su; Hui Wen Ng; Mao Shu; Heng Luo; Weigong Ge; Roger Perkins; Weida Tong; Huixiao Hong
Journal:  BMC Bioinformatics       Date:  2014-10-21       Impact factor: 3.169

  2 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.