| Literature DB >> 21414991 |
Jochen Weile1, Matthew Pocock, Simon J Cockell, Phillip Lord, James M Dewar, Eva-Maria Holstein, Darren Wilkinson, David Lydall, Jennifer Hallinan, Anil Wipat.
Abstract
MOTIVATION: The rise of high-throughput technologies in the post-genomic era has led to the production of large amounts of biological data. Many of these datasets are freely available on the Internet. Making optimal use of these data is a significant challenge for bioinformaticians. Various strategies for integrating data have been proposed to address this challenge. One of the most promising approaches is the development of semantically rich integrated datasets. Although well suited to computational manipulation, such integrated datasets are typically too large and complex for easy visualization and interactive exploration.Entities:
Mesh:
Year: 2011 PMID: 21414991 PMCID: PMC3077072 DOI: 10.1093/bioinformatics/btr134
Source DB: PubMed Journal: Bioinformatics ISSN: 1367-4803 Impact factor: 6.937
Data sources used in this work
| Data | Source | Version/date |
|---|---|---|
| Genome | SGD | 11/02/2010 |
| GO annotations | SGD | 11/02/2010 |
| Interactome | BioGRID | v2.0.61 |
| Regulatory network | NA | |
| Metabolic network | v1.0 | |
| Homology | BLAST (id. >85%) | NA |
NA = not applicable.
Fig. 1.Simplified schematic of the data structure underlying the Ondex Saccharomyces knowledge network. Circles represent concept types, and lines represent relations between them.
Fig. 2.Schema illustrating two examples for constructing associations based on metadata motifs. (A) The motif shown selects all paths from a polypeptide that is part of a transcription factor which regulates a gene that is transcribed to an mRNA which is translated into another polypeptide. (B) The motif from (A) is shown in context of the complete metadata structure. (C) A small example subnetwork to which the motif can be applied. Finding the motif (A) in the subnetwork, we identify three matching paths. Each element in these paths matches its corresponding element in the motif. (D) The view resulting from collapsing the paths identified in (C) according to motifs from (A) contains three edges; one for each matching path. (E) The motif selects all paths from a protein that modifies a process that produces a metabolite which is then consumed by another process back to another protein that modifies that process. Such a motif could be described as ‘precedence in a metabolic pathway’. (F) The motif from (E) is shown in context of the complete metadata structure. (G) A small example subnetwork to which the motif can be applied. Finding the motif (E) in the subnetwork, we identify two matching paths. Each element in these paths either matches or is a subtype of the corresponding element in the motif. (H) The views resulting from collapsing the paths identified in (G) according to motifs from (E) contains two edges; one for each matching path.
Motif definitions used in this work
| Association name | Motif |
|---|---|
Gene groups in the neighbourhood of BMH1 and BMH2 in the created view and their cardinalities
| Total | Cell cycle related No. (%) | Glucose metabolism related No. (%) | Histone related No. (%) | |
|---|---|---|---|---|
| Neighb. of | 68 | 15 (22.1) | 7 (10.3) | 4 (5.9) |
| Neighb. of | 79 | 15 (18.9) | 8 (10.1) | 4 (5.1) |
| Intersection of | 28 | 7 (25.0) | 3 (10.7) | 4 (14.3) |
| Union of | 119 | 23 (19.3) | 12 (10.1) | 4 (3.3) |
Fig. 3.Screenshots from OndexView, summarizing the hypotheses presented. Edges represent physical interactions (blue), homology (red), regulation (green) and epistasis (yellow). Nodes represent genes, some of which are weak (light red) and some strong (dark red) suppressors of cdc13-1. Genes which show lethal or slow growing phenotypes upon deletion are marked with dashed outlines, as knowledge of their epistatic behaviour can be expected to be incomplete.