| Literature DB >> 16414954 |
Hong-Yu Ou1, Ling-Ling Chen, James Lonnen, Roy R Chaudhuri, Ali Bin Thani, Rebecca Smith, Natalie J Garton, Jay Hinton, Mark Pallen, Michael R Barer, Kumar Rajakumar.
Abstract
We devised software tools to systematically investigate the contents and contexts of bacterial tRNA and tmRNA genes, which are known insertion hotspots for genomic islands (GIs). The strategy, based on MAUVE-facilitated multigenome comparisons, was used to examine 87 Escherichia coli MG1655 tRNA and tmRNA genes and their orthologues in E.coli EDL933, E.coli CFT073 and Shigella flexneri Sf301. Our approach identified 49 GIs occupying approximately 1.7 Mb that mapped to 18 tRNA genes, missing 2 but identifying a further 30 GIs as compared with Islander [Y. Mantri and K. P. Williams (2004), Nucleic Acids Res., 32, D55-D58]. All these GIs had many strain-specific CDS, anomalous GC contents and/or significant dinucleotide biases, consistent with foreign origins. Our analysis demonstrated marked conservation of sequences flanking both empty tRNA sites and tRNA-associated GIs across all four genomes. Remarkably, there were only 2 upstream and 5 downstream deletions adjacent to the 328 loci investigated. In silico PCR analysis based on conserved flanking regions was also used to interrogate hotspots in another eight completely or partially sequenced E.coli and Shigella genomes. The tools developed are ideal for the analysis of other bacterial species and will lead to in silico and experimental discovery of new genomic islands.Entities:
Mesh:
Substances:
Year: 2006 PMID: 16414954 PMCID: PMC1326021 DOI: 10.1093/nar/gnj005
Source DB: PubMed Journal: Nucleic Acids Res ISSN: 0305-1048 Impact factor: 16.971
Figure 1Flowchart depicting the tRNAcc high-throughput strategy developed and used to analyse the contents and contexts of tRNA genes in sequenced E.coli and Shigella genomes. Four stand-alone tools, indicated in bold italic font in the figure, were employed to identify islands (IdentifyIsland, TabulateIsland) and design primers (ExtractFlank, Primaclade) corresponding to the conserved upstream and downstream flanking regions of each tRNA site to be interrogated. See Table 1 for a summary of the programs features. In this study, four complete genomes were compared by the tRNAcc method: E.coli K-12 MG1655, E.coli UPEC CFT073, E.coli O157:H7 EDL933 and S.flexneri 2a Sf301. Four distinct genome subsets were analysed with the MG1655 genome being used as the reference template in each case. The numbers in the ovals above the word ‘tRNA’ indicate the number of tRNA genes still being considered at each stage in the analysis. The following abbreviations were used: UCB, upstream chromosomal block; DCB, downstream chromosomal block; GI, genomic island; UF, 2 kb upstream conserved flank; DF, 2 kb downstream conserved flank.
Figure 2Schematic representation of a range of hypothetical tRNA site configurations present in the four complete genomes (MG1655, CFT073, EDL933 and Sf301) (a–f). The conserved UF and DF regions flanking tRNA genes are shown as dark grey filled boxes. UF and DF boxes drawn below the line indicate inversions with respect to the reference template MG1655 (c and d). The UF and DF boxes shown in pale grey with a broken outline represent deletions with respect to MG1655 (e and f). Genomic islands, where present, are indicated as broken boxes to emphasize the relatively large size of these regions. Arrowheads shown below each sub-figure indicate the location and orientation of primers specific to the UF and DF regions. Hollow arrowheads indicate the absence of matching complementary sequence. The solid line between the arrowheads shown in (a) indicates a likely successful in vitro PCR amplification; while the dotted line in (b) indicates a successful e-PCR-based ‘amplification’ that would typically yield a product of size far in excess of that that could be generated through standard in vitro PCR. The numbers shown above each configuration after the colon symbol represent the number of examples observed in the four genomes tested based on the 87 MG1655 tRNA genes and the total complement of orthologues present in the other three genomes (Table 2). The numbers of examples observed in the five unpublished genomes (E.coli EAEC O42, EPEC E2348/69, ETEC E24377A, E.coli HS and S.sonnei 53G), with respect to the subset of 20 tRNA genes only (Supplementary Table S3), are shown in parentheses. The symbols shown alongside the drawings are used in Table 2 and Supplementary Table S4 to highlight tRNA loci affected by inversions and/or deletions. Examples of the various atypical configurations observed in the four genomes are shown to the right. The figure is not drawn to scale.
Stand-alone tools developed and used for high throughput analyses of the contents and contexts of tRNA genes in bacterial genomes
| Software tool | Description | Reference |
|---|---|---|
| Island identification | ||
| IdentifyIsland | Identify putative islands based on conserved flanking blocks recognized by the multiple aligner Mauve 1.2.2 ( | This work |
| TabulateIsland | Tabulate the identified islands following analysis of different subsets of genomes | This work |
| LocateHotspots | Locate proposed hotspots in non-annotated chromosomal sequences using BLASTN-based searches | This work |
| Primer design | ||
| ExtractFlank | Generate multi-FASTA files containing the upstream or downstream flanking regions for the identified islands | This work |
| Primaclade | Design conserved PCR primers for the upstream or downstream flanking regions found in multiple bacterial genomes being compared. This program is available at | ( |
| Island analysis | ||
| DNAnalyser | Calculate the GC content and dinucleotide bias of identified islands, and the negative cumulative GC profile of genomes | This work |
| GenomeSubtractor | High-throughput BLASTN-based comparison of CDS sequences against test genomes to identify strain-specific CDS based on the level of nucleotide similarity | This work |
aThese programs can also be used for the generic identification and preliminary characterization of putative genomic islands located at other user-specified hotspots and for the analysis of cognate flanking sequences.
Sizes of genomic islands identified by the tRNAcc method that map to tRNA sites in four sequenced E.coli and Shigella genomesa
aIsland sizes are shown to the nearest 0.1 kb. Predicted insertions at these loci of >1 kb in size are highlighted in bold type to indicate putative genomic islands. The sizes of the 21 tRNA-borne islands identified by Islander (21) are shown in square brackets. Details of the arrowhead symbols used are explained in Figure 2.
bThe identities of the 2 kb upstream flanking regions (UF) and the 2 kb downstream flanking regions (DF) across all the four genomes are calculated by the multiple alignment program ClustalW 1.82 (15). Note that genomes exhibiting deletions of particular flanking regions were excluded from the corresponding multiple sequence alignments. If the identity of the complete 2 kb flanking sequences was <90%, a highly conserved region within the UF or DF region was further investigated. The sizes and identities of these shorter highly conserved regions present within the 2 kb segments themselves are shown in parentheses.
cAs the metV gene was not annotated within the Sf301 genome, the Sf301 data shown relate to the sequence-identical metW gene that is immediately adjacent.
dA 4.3 kb DNA fragment with termini corresponding to the 3′ ends of the inversely orientated, sequence identical asnW and asnV genes was inverted in CFT073, with respect to the other three genomes. Consequently, the 54.4 kb island identified by Islander as being integrated into the gene annotated as asnW in CFT073 was missed by tRNAcc analysis. In this study, we have re-labelled this latter CFT073 gene as ‘asnV’ to maintain the synteny of this tRNA gene and its cognate DF sequence in all four genomes. Hence, the CFT073 NCBI annotated asnV gene was now known as ‘asnW’. The upstream identities for this locus were calculated using the 2 kb upstream flanking regions of asnV in MG1655, EDL933 and Sf301 and the 2 kb UF fragment corresponding to the newly termed asnW gene in CFT073. The downstream identities were computed based on the downstream flanking regions of asnV in MG1655, EDL933 and Sf301 and the conserved downstream flanking region lying distal to the 54.4 kb island in CFT073.
eOwing to a 6.9 kb deletion of the upstream flanking region in CFT073 and a 4.1 kb deletion of the upstream region in Sf301, with respect to MG1655 and EDL933, no island was identified in the argU sites of the four strains using the tRNAcc method. As Islander had identified a 21.3 kb argU-borne island in MG1655, tRNAcc was re-run using only the two genomes (MG1655 and EDL933), resulting in successful identification of the MG1655 island.
fThe secondary conserved downstream flanking region was inverted with respect to MG1655.
Summary of tRNA-borne genomic islands in four published E.coli and Shigella genomes identified with tRNAcc
| Strain | Annotated genome | Islands identified | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Chromosome size (kb) | No. of annotated CDS | No. of strain-specific CDS | No. of horizontally transferred CDS | No. of islands | Total size of islands (kb) | No. annotated island-borne CDS | No. of island-borne strain-specific CDS | No. of island-borne horizontally transferred CDS | |
| 4640 | 4242 | 234 | N/A | 12 | 184.9 | 188 | 89 (47.3%) | N/A | |
| 5231 | 5379 | 884 | N/A | 15 | 708.8 | 742 | 458 (61.7%) | N/A | |
| 5528 | 5324 | 916 | 593 | 14 | 554.1 | 621 | 380 (61.2%) | 245 (39.5%) | |
| 4607 | 4180 | 236 | N/A | 10 | 217.7 | 232 | 61 (26.3%) | N/A | |
aIdentification of CDS as strain-specific among the genomes compared is based on the level of nucleotide similarity. See Supplementary Data for details.
bThe horizontally transferred CDS in EDL933 were taken from Horizontal Gene Transfer Database (HGT-DB) () (28). The horizontally transferred CDS for CFT073, Sf301 and the updated version of the MG1655 genome are currently not available (N/A) in the HGT-DB.
cThe percentages of island-borne CDS that are defined as strain-specific in this study are shown in parentheses.
dThe percentages of island-borne CDS that are listed in the HGT-DB are shown in parentheses.
Figure 3The distribution of H-values corresponding to the island-borne CDS in E.coli K-12 MG1655, E.coli UPEC CFT073, E.coli O157:H7 EDL933 and S.flexneri 2a Sf301 identified by tRNAcc and/or Islander methods. This homology score had been proposed by Fukiya et al. (18) and reflected the degree of similarity between the matching reference genome sequence and the CDS itself in terms of the length of match and the degree of identity at a DNA level. See Supplementary Data for details. Red, green, blue, cyan and magenta bars represent total CDS, CFT073 CDS, EDL933 CDS, MG1655 CDS and Sf301 CDS, respectively. Note that each CDS in a given genome has three H-values that were obtained by BLASTN searches against the other three genomes in turn.
Figure 4Negative cumulative GC profile (23) highlighting the genomic context of islands identified by tRNAcc in E.coli O157:H7 EDL933. A sharp upward spike in the negative cumulative GC profile indicates a relatively sharp increase in GC content, whereas an abrupt fall indicates a relatively sharp decrease in GC content. The locations of tRNA-associated genomic islands are shown in green and the tRNA and tmRNA genes are represented as blue diamonds. Details of this plot are specified in the supplementary material.
Figure 5The four pheV-borne islands in MG1655 (a), CFT073 (b), Sf301 (c) and EDL933 (d) genomes identified by the tRNAcc method. The 9.1 kb island in MG1655, 127.9 kb island in CFT073 and 55.1 kb island in Sf301 are flanked by conserved upstream (UF) and downstream (DF) backbone segments. However, the DF region, common to the other three genomes, is absent in EDL933. Instead, the first instance of a conserved chromosomal block common to the other genomes occurs 23.5 kb downstream of the EDL933 pheV gene. This secondary conserved block has been designated as DF′. Matching 2 kb flanking regions are represented as connected blocks. In this study, genomic island-like regions were defined as anomalous segments between the 3′ end of tRNA genes and the 5′ end of the conserved downstream flank. Consequently, the tRNAcc-identified GIs in MG1655, CFT073 and Sf301 lay between the pheV and DF loci, while that in EDL933 was defined as the segment between the pheV gene and the proximal boundary of the DF′ conserved segment. The Islander-defined 104.6 and 46.7 kb islands at the pheV locus in CFT073 and Sf301, respectively, are shown as red lines flanked by DR sequences (red rectangles).
Summary of putative islands in the five unpublished E.coli and Shigella genomes identified by in silico tRIPa
| Strain | Unpublished genome | Islands identified | |||||
|---|---|---|---|---|---|---|---|
| Chromosome size (kb) | No. of CDS predicted | No. of strain-specific CDS | No. of islands | Total size of islands (kb) | No. of islands-borne predicted CDS | No. of island-borne strain-specific CDS | |
| 5242 | 4899 | 490 | 10 | 440.0 | 485 | 205 (42.3%) | |
| 5075 | 5313 | 606 | 10 | 208.0 | 241 | 115 (41.7%) | |
| 4980 | 4254 | 261 | 11 | 411.1 | 283 | 120 (42.2%) | |
| 4644 | 3989 | 146 | 9 | 140.6 | 79 | 29 (36.7%) | |
| 4989 | 5118 | 344 | 6 | 228.1 | 225 | 16 (7.1%) | |
aThe two unpublished genomes E.coli ETEC E24377A and E.coli O9 HS sequenced by TIGR were downloaded from NCBI. The other three unpublished chromosomal sequences of E.coli EAEC O42, E.coli EPEC E2348/69 and S.sonnei 53G were downloaded from the Sanger Institute (). The putative CDS of E.coli EPEC E2348/69 and S.sonnei 53G were identified using GLIMMER 2.13 (29).
bPutative islands were identified when in silico tRIP PCR amplicons were at least 1 kb greater in size than those corresponding to empty matching sites. These islands are highlighted in bold type in Table S3 in the supplementary materials.
cDetermination of strain-specific CDS among the genomes compared was based on the level of nucleotide similarity with respect to MG1655, CFT073, EDL933 and Sf301. See the supplementary materials for details.
dThe percentages of island-borne CDS that were defined as strain-specific in this study are shown in parentheses.