| Literature DB >> 31507629 |
Caio Rafael do Nascimento Santiago1,2, Renata de Almeida Barbosa Assis3, Leandro Marcio Moreira3,4, Luciano Antonio Digiampietri1,5.
Abstract
Genomi<span class="Chemical">cs research has produced an exponential amount of data. However, the genetic knowledge pertaining to certain phenotypic cha<span class="Chemical">racteristics is lacking. Also, a considerable part of these genomes have coding sequences (CDSs) with unknown functions, posing additional challenges to researchers. Phylogenetically close microorganisms share much of their CDSs, and certain phenotypes unique to a set of microorganisms may be the result of the genes found exclusively in those microorganisms. This study presents the GTACG framework, an easy-to-use tool for identifying in the subgroups of bacterial genomes whose microorganisms have common phenotypic characteristics, to find data that differentiates them from other associated genomes in a simple and fast way. The GTACG analysis is based on the formation of homologous CDS clusters from local alignments. The front-end is easy to use, and the installation packages have been developed to enable users lacking knowledge of programming languages or bioinformatics analyze high-throughput data using the tool. The validation of the GTACG framework has been carried out based on a case report involving a set of 161 genomes from the Xanthomonadaceae family, in which 19 families of orthologous proteins were found in 90% of the plant-associated genomes, allowing the identification of the proteins potentially associated with adaptation and virulence in plant tissue. The results show the potential use of GTACG in the search for new targets for molecular studies, and GTACG can be used as a research tool by biologists who lack advanced knowledge in the use of computational tools for bacterial comparative genomics.Entities:
Keywords: comparative genomics; gene families; orthologs; systems biology; user-friendly tools
Year: 2019 PMID: 31507629 PMCID: PMC6718126 DOI: 10.3389/fgene.2019.00725
Source DB: PubMed Journal: Front Genet ISSN: 1664-8021 Impact factor: 4.599
Figure 1Steps involved in constructing the GTACG pipeline separated into three stages of pre-processing: the identification of homologous genes, the comparison of genomes, and the visualization of data. To facilitate the visualization of the relationships that the data have in each of the activities, the arrows were colored as follows: in black is the general data on genomes; in blue is the data about groups of genomes; in red is the data on the sequences; in yellow are the graphical results for visualization.
Figure 2GTACG home screen. These results are divided into five sections: Settings, Filters, Statistics, 2D Plot, and Phylogeny. The first two sections are related to the subsequent family’s searches; the others are related to genome data. (A) The first allows the navigation between the different levels of clustering (homology, orthology, and domains). (B) The second allows filtering the presence/absence of the genomes or according to groups of genomes; this section also shows the number of genomes which are being filtered (label 1 in the figure) and the number of families after applying the filters (label 2 in the figure). (C) The third, Statistics, presents the graphs for the metrics related to families, sequences, and local alignments. (D) The fourth, 2D Plot, presents a bidimensional projection of the genomes. (E) Finally, Phylogeny presents the built phylogenies and customization options. Most sections fit users’ screen size.
Figure 3Screen containing the results of only one family. This screen contains four main sections followed by sections summarizing the family data in relation to the groups of genomes presented in the framework. (A) The first section has the sequence data and the data of their respective genomes; it is also possible to graphically visualize the position of each sequence in the genome as well as its vicinity. (B) In the next section, Phylogeny, it is possible to visualize, customize, and reconstruct (with different parameters) the phylogeny of the sequences. (C) The following section shows the alignment of all the sequences; it is possible to view, customize, and rebuild the alignment. (D) The fourth section presents the graph constructed to identify families, in which sequences are represented as vertices and local alignments as edges. The graph can be customized to highlight the alignments in accordance with some specific metrics. In this figure, the local alignments with identity equal to 100% are highlighted. (E) Finally, the last section summarizes the statistical data from each group of genomes with metrics about the number of genomes, dissimilarity, and MIST.”.
Figure 4Phylogenetic profiles established by GTACG from the input genomes. The phylogeny (A) was inferred using the binary vectors for each genome; the positions of the vector represent the families and are defined as 0 or 1, depending on the presence/absence of the genome in the respective family; the method of inference was the parsimony program (pars) for binary features in the Phylip package. The phylogeny (B) was constructed using the distance matrix (using the Euclidian distance) of the binary vectors referred to above; the inference method chosen was the neighbor-joining also available in the Phylip package. The phylogeny (C) was constructed using a supertree that summarizes the collection of all the phylogeny constructed for the families; the tree of each family was obtained using the Clustal Omega to make the alignments and after that the FastTree produce the trees; the supertree method was the Quartet Fit algorithm with Nearest Neighbour Interchange available in the Clann.
Comparison of the main functionalities of some comparative genomics frameworks.
| GTACG | BPGA | PanX | PGAT | PanGP | PGAP | Panseq | ITEP | Get Homologues | |
|---|---|---|---|---|---|---|---|---|---|
| Identification of phenotype-specific genes – list | X | X | X | X | |||||
| Identification of phenotype-specific genes – metrics | X | ||||||||
| Distribution of core, accessory and unique genes | X | X | X | ||||||
| Pangenome profile analysis | X | X | X | X | X | ||||
| Size of core and pan-genome | X | X | X | X | X | X | |||
| Extraction of core, accessory and unique genes’ sequence | X | X | X | ||||||
| Evolutionary analysis | X | X | X | X | X | X | X | ||
| Protein/gene clustering | X | X | X | X | X | X | X | X | |
| Multilevel perspective of the genes | X | X | X | X | |||||
| Input data from user | X | X | X | X | X | X | |||
| Easy to share results | X | X | X | ||||||
| Integration with roary scripts | X | ||||||||
| Data preparation | C | C | C | N/A | G | C | C | C | C |
| User interface | W | GO | W | W | GO | GO | GO | GO | GO |
| References |
|
|
|
|
|
|
|
|
Data preparation: C, Command line; G, Graphical interface.
User interface: W, Website; GO, Graphical output.
Figure 5Relative runtime for GTACG’s main tasks with different datasets of Xanthomonas genomes. These results were obtained using a computer with an Intel(R) Xeon(R) CPU E5-2620. This computer has 24 cores, but only 20 of them were used. As Blast's alignments correspond to the majority of the consumed time, section (A) present the time spent excluding the time spent with Blast, while section (B) present the time including Blast.
Figure 6Runtime with different datasets of Xanthomonas genomes. The results show the execution time growth of the GTACG’s main tasks, according to the number of genomes in the datasets.
Figure 7Family of orthologous sequences in which all sequences from plant-associated genomes are isolated from the other genomes.
Characterization of the 18 protein families exclusively identified in genomes of bacteria associated with plants.
| Function | Gene name | Ref. Locus Tag | # Genomes | # Paralogs | Pathway | SP | Refs |
|---|---|---|---|---|---|---|---|
| Conserved hypothetical protein (putative lipase) |
| XAC0501 | 134 | 27 | Lipid metabolism | N |
|
| Peptidase M16 family/Zinc protease/Insulinase family protein | — | XAC0609 | 138 | 1 | Peptidases | Y |
|
| Low molecular weight heat shock protein/Molecular chaperone |
| XAC1151 | 138 | 1 | Chaperones and folding catalysis | N |
|
| Cytochrome O ubiquinol oxidase subunit IV |
| XAC1261 | 138 | 2 | Oxidative phosphoryla-tion | N |
|
| Conserved hypothetical protein | — | XAC2544 | 137 | 2 | Unknown function | Y | — |
| Predicted 4-hydroxyproline dipeptidase/Xaa-Pro aminopeptidase |
| XAC2545 | 138 | 1 | Metallo peptidases | N | — |
| Alpha-L-fucosidase |
| XAC3072 | 138 | 1 | N-glycan metabolism | Y |
|
| Hypothetical protein (putative glycosyl-hydrolase) |
| XAC3073 | 138 | 1 | N-glycan metabolism | Y |
|
| Beta-hexosaminidase/Beta-N-acetylglucosaminidase |
| XAC3074 | 138 | 1 | N-glycan metabolism | Y |
|
| Beta-mannosidase |
| XAC3075 | 138 | 3 | N-glycan metabolism | Y |
|
| Beta-glucosidase-related glycosidases/Gluca-beta-glucosidase |
| XAC3076 | 138 | 2 | N-glycan metabolism | Y |
|
| Hypothetical protein (putative glycosyl-hydrolase) |
| XAC3082 | 138 | 4 | N-glycan metabolism | Y |
|
| Alpha-1,2-mannosidase |
| XAC3083 | 138 | 1 | N-glycan metabolism | N |
|
| Beta-galactosidase |
| XAC3084 | 138 | 1 | N-glycan metabolism | N |
|
| Cytoplasmic copper homeostasis protein CutC |
| XAC3091 | 138 | 2 | Copper metabolism | N | — |
| 3-isopropylmalate dehydrogenase/Isocitrate dehydrogenase |
| XAC3456 | 134 | 1 | Leucine biosynthesis | N |
|
| Integral membrane protein | — | XAC4076 | 134 | 1 | Unknown function | N | — |
| N-acetylglucosamine-regulated/TonB-dependent receptor |
| XAC4131/3071 | 138 | 10 | TonB receptors/N-glycan metabolism | Y |
|
| Conserved hypothetical protein | — | XAC4164 | 137 | 1 | Unknown function | Y |
|
SP, signal peptide; Y, yes; N, no.
Figure 8Phylogenetic analysis of 8 out of 19 protein families identified only among the genomes associated with the plants belonging to the family Xanthomonadaceae. The identification of circles, colors, and sizes is not provided by the tool; they have been inserted in this context only to facilitate the description of the identifiers. It is possible to observe a pattern in the topology of the phylogenies of the hydrolases, always with larger branches for organisms of the genus Xylella and Xanthomonas translucens, X. sacchari, and X. albilineans. (A) alpha-L-fucosidase family. (B) beta-galactosidase family. (C) beta-glucosidase-related glycosidases family. (D) glycosyl hydrolase family. (E) beta-N-acetylglucosaminidase family. (F) 4-hydroxyproline dipeptidase family. (G) beta-mannosidase family. (H) alpha-mannosidase family.
Figure 9Identification of the genes related with plant N-glycan degradation. (A) N-Glycan metabolism gene cluster in Xac306 genome. Red – Genes identified as exclusive of plant-associated genomes. The numbers 1 to 10 identify all genes related to N-glycan degradation. a – Non-related to N-glycan degradation. (B) Model of plant N-glycan structure. The numbers 1 to 10 identify the catalytic site of the proteins coded by the genes described in (A). Asn – Asparagine residue. Ser/Thr – Serine and threonine residues. X – Any residue.