| Literature DB >> 17937820 |
William A McLaughlin1, Ken Chen, Tingjun Hou, Wei Wang.
Abstract
BACKGROUND: Protein domains coordinate to perform multifaceted cellular functions, and domain combinations serve as the functional building blocks of the cell. The available methods to identify functional domain combinations are limited in their scope, e.g. to the identification of combinations falling within individual proteins or within specific regions in a translated genome. Further effort is needed to identify groups of domains that span across two or more proteins and are linked by a cooperative function. Such functional domain combinations can be useful for protein annotation.Entities:
Mesh:
Substances:
Year: 2007 PMID: 17937820 PMCID: PMC2151957 DOI: 10.1186/1471-2105-8-390
Source DB: PubMed Journal: BMC Bioinformatics ISSN: 1471-2105 Impact factor: 3.169
Figure 1An illustration of the derivation of a DASSEM unit and its functional annotation. For the group of proteins shown, there are three prevalent domain compositions within individual proteins (circled). The domain fusions of these prevalent domain compositions were used to link three domains and to create a DASSEM unit that contains the fork head (FH) domain, the fork head associated (FHA) domain and the kinase domain. The overall function of the DASSEM unit was obtained by finding the GO terms that were enriched across the proteins associated with the unit. The GO term enrichment indicated that the unit participates in the cell cycle. Schematics of domains and proteins were taken from the Pfam database [89].
Six example domain assembly units and their corresponding annotations.
| FHA domain (PF00498) | P-cell cycle 4.32*10-06 |
| Fork head domain (PF00250) | C-nucleus 2.12*10-02 |
| Protein kinase domain (PF00069) | F-transcription factor activity 4.22*10-05 |
| Fungal Zn(2)-Cys(6) domain (PF00172) | P-regulation of transcription: 8.33*10-17 |
| Fungal transcription factor domain (PF04082) | C-nucleus, 5.48*10-12 |
| Gal4-like dimerization domain (PF03902) | F-transcription regulator, 2.24*10-26 |
| PAS fold domain (PF00989) | |
| ABC transporter domain (PF00005) | P-transport, 3.02*10-05 |
| ABC transporter region 1 domain (PF00664) | C-membrane, 4.15*10-05 |
| ABC transporter region 2 domain (PF06472) | F-ATPase activity, 5.68*10-23 |
| ABC-2 type transporter (PF01061) | |
| Metal-binding domain in RNase L domain (PF04068) | |
| 4Fe-4S binding domain (PF00037) | |
| GTPase of unknown function domain (PF01926) | P-ribosome-nucleus export, 1.51*10-02 |
| DUF933 domain (PF06071) | C-mitochondrial inner membrane, 3.27*10-02 |
| TGS domain (PF02824) | F-GTP binding, 9.86*10-09 |
| GTP1/OBG domain (PF01018) | |
| Helicase conserved C-terminal domain (PF00271) | P-chromosome organization, 5.41*10-06 |
| SNF2 family N-terminal domain (PF00176) | C-chromatin remodeling complex, 1.77*10-03 |
| HSA (PF07529) | F-ATPase activity 2.04*10-07 |
| Chromatin organisation modifier domain (PF00385) | |
| Elongation factor Tu GTP binding domain (PF00009) | P-translation factor activity 1.08*10-20 |
| Elongation factor Tu domain 2 (PF03144) | C-ribosome 3.71*10-04 |
| Elongation factor G C-terminus (PF00679) | F-translation 1.33*10-09 |
| Elongation factor G, domain IV (PF03764) | |
| Elongation factor Tu C-terminal domain (PF03143) | |
| GTP-binding protein LepA C-terminus (PF06421) | |
| Translation initiation factor IF-2, N-term. (PF04760) |
Each domain assembly unit or DASSEM unit is described as a list of Pfam domains and its GO term annotation. The terms which have the most significant p-values for the GO categories of biological process (P), cellular component (C), and molecular function (F), are shown for each unit. The full lists of enriched GO terms are provided in a companion website [33].
Figure 2An illustration of the utilization of DASSEM units within a transcription module. The example transcription module is involved in the process of amino acid biosynthesis, and the DASSEM units contribute to necessary auxiliary processes. The terms listed are from the GO term "biological process" category. M- a transcription module involved in amino acid biosynthesis, 1- a unit involved in aromatic carbon metabolism, 2- a unit involved in sulfate assimilation, 3- a unit involved in serine biosynthesis, 4- a unit involved in ethanol metabolism, and 5- a unit involved in amino acid derivative metabolism. The equation for the overlap score is given along with the overlap scores for the DASSEM units in the example.
Figure 3Plots of the overlap scores of doma in content of DASSEM units with that of transcription modules. Overlap scores of the DASSEM units with transcription modules are given for before (black) and after (white) randomization of the domains in the modules. The highest overlap scores, where one DASSEM unit was paired with each transcription module, are shown in panel A. The overlap scores were also calculated when a collection of five DASSEM used were paired with each transcription module. These are shown are shown in panel B. The plots indicate the DASSEM units were utilized in transcription modules, based their overlap of domain content.
Figure 4Plot of the overlap scores of protein content of DASSEM units with that of transcription modules. The overlap scores of the DASSEM units with transcription modules are given for before (black) and after (white) randomization of the proteins in the modules. The highest overlap scores, where one DASSEM unit was paired with each transcription module, are shown. Student t-tests were used to compare the average overlap scores of the DASSEM unit with the original versus the randomized modules. The p-value of a Student's t-test that compared the two averages was significant at 0.004. Subsequent t-tests compared the second, third, fourth, and fifth highest overlap scores between the DASSEM units with the original or randomized modules. Their p-values were also significant at 0.0029, 0.0033, 0.0036, and 0.046 respectively.
Figure 5The distributions of GO term p-values and hierarchy levels for the DASSEM units, the transcription modules, and the random protein sets. The p-values of GO term enrichments for the DASSEM units (black), the transcription modules (gray), and the random sets of proteins (white) are shown in panel A. Since the range of p-values was large, the second logarithm, i.e. the logarithm of the absolute value of the first logarithm, was plotted for ease of visualization. The plot indicates that the number of GO terms and the values of the p-values were similar between the transcription modules and the DASSEM units. In contrast, there were much less terms associated with random sets of proteins, and the p-values of these terms were less significant. Panel B shows the levels of the GO terms within the GO hierarchy. For the transcription modules and the DASSEM units the depths of the GO term levels were similar. In contrast, the terms for the random protein sets were distributed at the higher GO levels where the terms are less specific.
Figure 6A schematic of how the DASSEM units were used to annotate proteins of unknown function. Domains are represented as colored blocks. If a protein of unknown function contains some of the domains of a DASSEM unit then it is likely to have all or part of the unit's function.