BACKGROUND: This article addresses the problem of interoperation of heterogeneous bioinformatics databases. RESULTS: We introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research. CONCLUSION: BioWarehouse embodies significant progress on the database integration problem for bioinformatics.
BACKGROUND: This article addresses the problem of interoperation of heterogeneous bioinformatics databases. RESULTS: We introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research. CONCLUSION: BioWarehouse embodies significant progress on the database integration problem for bioinformatics.
Authors: Peter D Karp; Monica Riley; Milton Saier; Ian T Paulsen; Julio Collado-Vides; Suzanne M Paley; Alida Pellegrini-Toole; César Bonavides; Socorro Gama-Castro Journal: Nucleic Acids Res Date: 2002-01-01 Impact factor: 16.971
Authors: Cynthia J Krieger; Peifen Zhang; Lukas A Mueller; Alfred Wang; Suzanne Paley; Martha Arnaud; John Pick; Seung Y Rhee; Peter D Karp Journal: Nucleic Acids Res Date: 2004-01-01 Impact factor: 16.971
Authors: Daniel Segrè; Jeremy Zucker; Jeremy Katz; Xiaoxia Lin; Patrik D'haeseleer; Wayne P Rindone; Peter Kharchenko; Dat H Nguyen; Matthew A Wright; George M Church Journal: OMICS Date: 2003
Authors: Oliver Ruebenacker; Ion I Moraru; James C Schaff; Michael L Blinov Journal: Proceedings (IEEE Int Conf Bioinformatics Biomed) Date: 2007-11-02
Authors: Peter D Karp; Mario Latendresse; Suzanne M Paley; Markus Krummenacker; Quang D Ong; Richard Billington; Anamika Kothari; Daniel Weaver; Thomas Lee; Pallavi Subhraveti; Aaron Spaulding; Carol Fulcher; Ingrid M Keseler; Ron Caspi Journal: Brief Bioinform Date: 2015-10-10 Impact factor: 11.622
Authors: Paolo Pannarale; Domenico Catalano; Giorgio De Caro; Giorgio Grillo; Pietro Leo; Graziano Pappadà; Francesco Rubino; Gaetano Scioscia; Flavio Licciulli Journal: BMC Bioinformatics Date: 2012-03-28 Impact factor: 3.169
Authors: Pragya A Dang; Mannudeep K Kalra; Michael A Blake; Thomas J Schultz; Markus Stout; Elkan F Halpern; Keith J Dreyer Journal: J Digit Imaging Date: 2008-06-10 Impact factor: 4.056