Literature DB >> 16556315

BioWarehouse: a bioinformatics database warehouse toolkit.

Thomas J Lee1, Yannick Pouliot, Valerie Wagner, Priyanka Gupta, David W J Stringer-Calvert, Jessica D Tenenbaum, Peter D Karp.   

Abstract

BACKGROUND: This article addresses the problem of interoperation of heterogeneous bioinformatics databases.
RESULTS: We introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research.
CONCLUSION: BioWarehouse embodies significant progress on the database integration problem for bioinformatics.

Entities:  

Mesh:

Substances:

Year:  2006        PMID: 16556315      PMCID: PMC1444936          DOI: 10.1186/1471-2105-7-170

Source DB:  PubMed          Journal:  BMC Bioinformatics        ISSN: 1471-2105            Impact factor:   3.169


  23 in total

1.  An ontology for biological function based on molecular interactions.

Authors:  P D Karp
Journal:  Bioinformatics       Date:  2000-03       Impact factor: 6.937

2.  TAMBIS: transparent access to multiple bioinformatics information sources.

Authors:  R Stevens; P Baker; S Bechhofer; G Ng; A Jacoby; N W Paton; C A Goble; A Brass
Journal:  Bioinformatics       Date:  2000-02       Impact factor: 6.937

3.  The EcoCyc Database.

Authors:  Peter D Karp; Monica Riley; Milton Saier; Ian T Paulsen; Julio Collado-Vides; Suzanne M Paley; Alida Pellegrini-Toole; César Bonavides; Socorro Gama-Castro
Journal:  Nucleic Acids Res       Date:  2002-01-01       Impact factor: 16.971

4.  The Molecular Biology Database Collection: 2004 update.

Authors:  Michael Y Galperin
Journal:  Nucleic Acids Res       Date:  2004-01-01       Impact factor: 16.971

5.  MetaCyc: a multiorganism database of metabolic pathways and enzymes.

Authors:  Cynthia J Krieger; Peifen Zhang; Lukas A Mueller; Alfred Wang; Suzanne Paley; Martha Arnaud; John Pick; Seung Y Rhee; Peter D Karp
Journal:  Nucleic Acids Res       Date:  2004-01-01       Impact factor: 16.971

6.  BioSPICE: access to the most current computational tools for biologists.

Authors:  Thomas D Garvey; Patrick Lincoln; Charles John Pedersen; David Martin; Mark Johnson
Journal:  OMICS       Date:  2003

7.  EnsMart: a generic system for fast and flexible access to biological data.

Authors:  Arek Kasprzyk; Damian Keefe; Damian Smedley; Darin London; William Spooner; Craig Melsopp; Martin Hammond; Philippe Rocca-Serra; Tony Cox; Ewan Birney
Journal:  Genome Res       Date:  2004-01       Impact factor: 9.043

8.  From annotated genomes to metabolic flux models and kinetic parameter fitting.

Authors:  Daniel Segrè; Jeremy Zucker; Jeremy Katz; Xiaoxia Lin; Patrik D'haeseleer; Wayne P Rindone; Peter Kharchenko; Dat H Nguyen; Matthew A Wright; George M Church
Journal:  OMICS       Date:  2003

9.  A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

Authors:  Michelle L Green; Peter D Karp
Journal:  BMC Bioinformatics       Date:  2004-06-09       Impact factor: 3.169

10.  Call for an enzyme genomics initiative.

Authors:  Peter D Karp
Journal:  Genome Biol       Date:  2004-07-30       Impact factor: 13.583

View more
  43 in total

1.  Flexible network reconstruction from relational databases with Cytoscape and CytoSQL.

Authors:  Kris Laukens; Jens Hollunder; Thanh Hai Dang; Geert De Jaeger; Martin Kuiper; Erwin Witters; Alain Verschoren; Koenraad Van Leemput
Journal:  BMC Bioinformatics       Date:  2010-07-01       Impact factor: 3.169

2.  Kinetic Modeling using BioPAX ontology.

Authors:  Oliver Ruebenacker; Ion I Moraru; James C Schaff; Michael L Blinov
Journal:  Proceedings (IEEE Int Conf Bioinformatics Biomed)       Date:  2007-11-02

3.  Pathway Tools version 19.0 update: software for pathway/genome informatics and systems biology.

Authors:  Peter D Karp; Mario Latendresse; Suzanne M Paley; Markus Krummenacker; Quang D Ong; Richard Billington; Anamika Kothari; Daniel Weaver; Thomas Lee; Pallavi Subhraveti; Aaron Spaulding; Carol Fulcher; Ingrid M Keseler; Ron Caspi
Journal:  Brief Bioinform       Date:  2015-10-10       Impact factor: 11.622

Review 4.  Network-based approaches in drug discovery and early development.

Authors:  J M Harrold; M Ramanathan; D E Mager
Journal:  Clin Pharmacol Ther       Date:  2013-09-11       Impact factor: 6.875

5.  Using web ontology language to integrate heterogeneous databases in the neurosciences.

Authors:  Hugo Y K Lam; Luis Marenco; Gordon M Shepherd; Perry L Miller; Kei-Hoi Cheung
Journal:  AMIA Annu Symp Proc       Date:  2006

6.  GIDL: a rule based expert system for GenBank Intelligent Data Loading into the Molecular Biodiversity Database.

Authors:  Paolo Pannarale; Domenico Catalano; Giorgio De Caro; Giorgio Grillo; Pietro Leo; Graziano Pappadà; Francesco Rubino; Gaetano Scioscia; Flavio Licciulli
Journal:  BMC Bioinformatics       Date:  2012-03-28       Impact factor: 3.169

Review 7.  Network integration and graph analysis in mammalian molecular systems biology.

Authors:  A Ma'ayan
Journal:  IET Syst Biol       Date:  2008-09       Impact factor: 1.615

8.  bioDBnet: the biological database network.

Authors:  Uma Mudunuri; Anney Che; Ming Yi; Robert M Stephens
Journal:  Bioinformatics       Date:  2009-01-07       Impact factor: 6.937

9.  Use of Radcube for extraction of finding trends in a large radiology practice.

Authors:  Pragya A Dang; Mannudeep K Kalra; Michael A Blake; Thomas J Schultz; Markus Stout; Elkan F Halpern; Keith J Dreyer
Journal:  J Digit Imaging       Date:  2008-06-10       Impact factor: 4.056

10.  Bringing Web 2.0 to bioinformatics.

Authors:  Zhang Zhang; Kei-Hoi Cheung; Jeffrey P Townsend
Journal:  Brief Bioinform       Date:  2008-10-08       Impact factor: 11.622

View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.