Literature DB >> 25792768

A SELEX study of the DNA-binding specificity of archaeal FFRPs: 2. FL4 (pot1613368).

Katsushi Yokoyama1, Michiko Ihara1, Sonomi Ebihara1, Masashi Suzuki1.   

Abstract

The DNA-binding specificity of a transcription factor, the FFRP FL4 (pot1613368) from Pyrococcus sp. OT3, was studied. Using SELEX (systematic evolution of ligands by exponential environment) experiments, from a set of fragments, ∼150 bps, of the genomic DNA of P. OT3, seven were selected as containing binding sites. Thirteen bases identified as shared by the seven selected fragments with the least mismatches, 2.71 on average, was ATGAA AAAGTCAT. This sequence was closely related with another sequence, ATGAA[AAA/TTT]TTCAT, in the 5-3-5 arrangement, i.e. NANBNCNDNE [AAA/TTT]NENDNCNBNA , where, e.g. NA was the base complementary to NA. The average number of mismatches found between this sequence and the seven fragments was 3.14. A sequence, TTGAA ATT TACAA, resembling the sequence ATGAA[AAA/TTT]TTCAT and also another 5-3-5 sequence, TTGAA[AAA/TTT]TTCAA, was found upstream of the fl4 gene, which is potentially recognized by FL4 for auto-regulation. Thus it is likely that an ideal binding-site of FL4 is ATGAA[AAA/TTT]TTCAT or TTGAA[AAA/TTT]TTCAA. In this abstract, the sequences were highlighted in Italic at 3, and with bold characters at 5 and 5. When two sequences compared were the same at some positions, there they were underlined.

Entities:  

Keywords:  Archaea; AsnC; DNA-protein interaction; Lrp; hyper-thermophile; transcription factor

Year:  2006        PMID: 25792768      PMCID: PMC4322925          DOI: 10.2183/pjab.82.33

Source DB:  PubMed          Journal:  Proc Jpn Acad Ser B Phys Biol Sci        ISSN: 0386-2208            Impact factor:   3.493


Introduction

We have been studying the structure and function of feast/famine regulatory proteins (FFRPs), regulating transcription of many genes in archaea and eubacteria.[1)–25)] The nucleotide sequences of DNA sites bound by various FFRPs are summarized into the same form, N[AAA/TTT]N, where, e.g. N is the base complementary to N[1),4),18),22)]: here referred to as the 5-3-5 arrangement. In the reverse direction, several transcription factors experimentally identified as binding DNA sequences in this arrangement have been re-identified as FFRPs by analyzing their amino acid sequences carefully.[4)] Thus the 5-3-5 arrangement appears to be uniquely associated with FFRPs. When 13 bps in the 5-3-5 arrangement are positioned by overlapping onto a TATA box or its downstream, transcription of the gene will be repressed by binding by an FFRP.[1),9)] While, when such 13 bps are positioned immediately upstream of a TATA box with an insertion of ∼4 or ∼15 bps (i.e. 4 plus 10.5), binding of an FFRP will activate transcription, possibly through its interaction with the TATA-binding protein (TBP), thereby recruiting TBP to the TATA box.[1),5),9)] In this paper, results obtained by SELEX (systematic evolution of ligands by exponential environment)[26)] experiments are analyzed in order to determine the DNA-binding specificity of an FFRP, FL4 (pot1613368) from a hyper-thermophilic archaeon, Pyrococcus sp. OT3. This FFRP is one of sixteen FFRPs we have identified as coded in the genome[27)] of this organism (Table I). The experiments which we report in this paper were, in fact, carried out several years ago. Yet many possible consensus sequences can be deduced from seven fragments of ∼270 bps each selected, and so by statistical analyses alone we were unable to pinpoint a short consensus sequence uniquely. Only recently by assuming a 5-3-5 arrangement for FL4-binding sites, we have come to a conclusion.
Table I.

Transcription factors, FFRPs, identified as coded in the genome of Pyrococcus sp. OT3

IDFuller ID*Orthologues from other speciesCrystal 3DEM analysisLigandBinding DNA-sequences
DM1pot1216151none++isoleucineN.A.
DM2pot0300646noneN.D.N.E.N.D.N.A.
DM3pot0175330noneN.D.N.E.N.D.N.A.
FL1pot0828564noneN.D.N.E.N.D.N.D.
FL2pot0836696noneN.D.N.E.N.D.N.D.
FL3pot0868477noneN.D.N.E.N.D.N.D.
FL4pot1613368noneN.D.N.E.N.D.ATGAA**
FL5pot1664679noneN.D.N.E.N.D.N.D.
FL6pot1735659noneN.D.N.E.N.D.N.D.
FL7pot0008824noneN.D.N.E.N.D.N.D.
FL8pot0123002noneN.D.N.E.N.D.N.D.
FL9pot0301583noneN.D.N.E.N.D.N.D.
FL10pot0377090LrpA from P. f.+N.D.TTCG[2)]
FL11pot0434017none++(glutamine)TGAAA[6)]
FL12pot0258936Phr from P. f.N.D.TAACC[4)]
FL13pot0846474TrmB from T. l.trehalose/maltoseATACT[4)]

N.A.: not applicable since these do not have a DBD. N.D.: not determined. N.E.: not examined. P. f.: Pyrococcus furiosus. T. l.: Thermococcus litralis.

pot (Pyrococcus sp. OT3) followed by stop codon positions in the genome (http://www.aist.go.jp/RIODB/archaic).

this study.

Transcription factors, FFRPs, identified as coded in the genome of Pyrococcus sp. OT3 N.A.: not applicable since these do not have a DBD. N.D.: not determined. N.E.: not examined. P. f.: Pyrococcus furiosus. T. l.: Thermococcus litralis. pot (Pyrococcus sp. OT3) followed by stop codon positions in the genome (http://www.aist.go.jp/RIODB/archaic). this study.

Materials and methods

Protein purification

The gene of the FL4 protein from Pyrococcus sp. OT3 was amplified by the polymerase chain reaction (PCR),[28)] and cloned into the pET28 expression vector. A protein expressed using this vector has a His-tag[29)] at its N-terminus. The E. coli strain BL21(DE3)/plysE was transformed with the vector, and the gene fl4 was expressed, using an inducer, isopropyl β-D-thiogalactopyranoside (IPTG). From a culture, 2 l, E. coli cells were collected by centrifugation at 9,000 × g for 10 min at 4 °C, and suspended into 15 ml of PBS buffer (0.4 mM Na2HPO4 and 0.18 mM KH2PO4, adjusted by HCl to pH, 7.4, containing 13.7 mM NaCl and 0.27 mM KCl) containing 1% Triton X100. The supernatant was sonicated twice for 30 sec each, and kept at 75 °C for 10 min, while mixed gently by pipeting. After centrifugation at 27,000 × g for 10 min at 25 °C, 4 M (NH4)2SO4, 30 ml, was added to the super-natant and kept at room temperature for 1 hr. After centrifugation at 9,000 × g for 10 min at 25 °C the sediment was dissolved into 2 ml of buffer A, i.e. 20 mM HEPES (adjusted to pH, 7.4, using KOH) containing 10 mM MgSO4, 1 mM DTT, 1 mM EDTA, 50 mM NaCl, and 5% glycerol, and kept at 85 °C for 10 min. After centrifugation at 9,000 × g for 10 min at 25 °C, the supernatant was dialyzed against 400 ml of buffer A for 2 hrs at room temperature. After centrifugation at 9,000 × g for 10 min at 25 °C, the protein precipitated was dissolved into 1 ml of buffer A containing 1 M urea, and dialyzed against 250 ml of buffer B, 14.3 mM HEPES (adjusted to pH, 7.4, using KOH) containing 7 mM MgSO4, 0.7 mM DTT, 0.7 mM EDTA, 270 mM NaCl, and 27% glycerol, for 2 hrs at room temperature. This process of dialysis was carried out once more overnight. The protein solution was centrifuged at 27,000 × g for 10 min at 25 °C, and filtered through a membrane (the pore size of 0.45 µm, NALGEN, Rochester) to remove large contaminants. The purified protein formed a single band in an SDS polyacrylamide gel after electrophoresis. Its binding to a column, Ni-NTA spin (QIAGEN, Hilden, Germany), using the His-tag added to its N-terminus, was confirmed.

Preparation of genomic DNA fragments

The genomic DNA molecule of P. OT3 was treated with restriction enzymes, AfaI, AluI, HaeIII, TthHB, HinfI, MseI, Sau3AI, ApaI and MboII, respectively. A mixture of these fragments were subjected to electrophoresis using a gel containing 1.5% SeaPlaque agarose (TaKaRa, Tokyo). From the gel fragments of the size, 100–1,000 bps, were recovered using SUPREC-01 (TaKaRa, Tokyo). Using the DNA blunting kit (TaKaRa, Tokyo), the fragments were ligated to the HincII site of the pBluescript plasmid, pre-treated with bacterial alkaline phosphatase. The plasmid containing various DNA fragments was introduced into an E. coli strain, XL-1-Blue. The E. coli cells were grown on LB plates containing ampicillin, 100 µg/ml, 5-bromo-4-chloro-3-indolyl-β-D-galactoside (X-gal), 130 µg/ml, and IPTG, 1 mM, overnight at 37 °C. A mixture of the plasmid DNAs amplified (see Results) was used as a DNA library in the following round of SELEX experiments.

SELEX Protocol

The DNA library, 5.0 µg of the pBluescript plasmids, FL4, 1.0 µg, and nickel-coated silica beads, 10 µl of a suspension, 250 µl, of materials obtained from a column, Ni-NTA spin (QIAGEN, Hilden, Germany), were mixed into PBS buffer, 100 µl, containing 20% glycerol, 10 mM imidazole, 1.0 µg poly dI/dC, 5 mM β-mercaptoethanol, and varying concentrations of MgSO4, NaCl, and KCl (Table II). The solution was incubated for 15 min at a temperature, 60, 80 or 95 °C. As a negative control, 0.2 µg of the library DNA was used instead of 5.0 µg (see Results).
Table II.

Conditions and efficiencies of SELEX experiments

DNA (µg)FL4 (µg)MgSO4 (mM)KCl (mM)NaCl (mM)°CNo. whiteNo. bluePrimary W/B*Secondary W/B**
Optimization 1
5.01.05150150605908180.721.07
5.01.025150150605637800.721.07
5.01.055150150602272081.091.63
5.01.01051501506017101.702.53
5.01.05350150604839870.490.73
5.01.05150350603315480.600.89
0.21.05150150601161720.671.00

Optimization 2
5.01.055150450602993250.921.61
5.01.0551505506084561.502.63
5.01.0551507506083471.773.10
5.01.0551501000602331211.933.39
0.21.05150150602364110.571.00

Optimization 3
5.01.055150450601152240.511.65
5.01.0551504508035810680.341.10
5.01.055150450953407880.431.39
5.01.055150100060721010.712.29
5.01.055150100080761450.521.68
5.01.055150100095841930.441.42
0.21.05150150601886150.311.00

1st round
5.01.0501501000603202701.192.28
0.21.0501501000603176100.521.00
5.00501501000604469620.460.88

2nd round
5.01.0501501000601842470.741.34
0.21.0501501000602564660.551.00
5.0050150100060115243280.270.49

the ratio of No. white to No. blue.

the W/B ratio relative to another W/B observed using 0.2 µg DNA.

Conditions and efficiencies of SELEX experiments the ratio of No. white to No. blue. the W/B ratio relative to another W/B observed using 0.2 µg DNA. After the incubation, silica beads were collected by centrifugation at 800 × g at 25 °C, and, after 500 µl of PBS buffer was added, centrifuged again. This washing process was repeated five times. Then, 200 µl of Tris-EDTA buffer, i.e. 10 mM Tris-HCl buffer (pH = 8.0) containing 0.1 mM EDTA, were added and mixed well. After centrifugation at 800 × g at 25 °C, phenol, 200 µl, was added, and plasmid DNAs were isolated by ethanol precipitation. The plasmid DNAs were suspended into 20 µl of Tris-EDTA buffer. Using 0.5 µl of this plasmid solution E. coli cells XL1-Blue were transformed.

Results

Strategy for optimizing the SELEX protocol

Host E. coli cells transformed with the original pBluescript plasmid are expected to form colonies in a bluish color. This plasmid carries the lacZ gene. Its product, β-galactosidase, catalyzes the substrate X-gal, present in the LB plate, thereby producing this color. When a DNA fragment is inserted into the lacZ gene, β-galactosidase will not be expressed in its original form (Fig. 1a, P2 and P3), yielding the original whitish color of E. coli cells. Thus, the SELEX protocol was optimized, so that the highest ratio of white to blue (W/B) was obtained, and so that the number of white colonies was reasonably high. Here, experiments carried out in the presence of plasmid DNA, 0.2 µg, were considered as negative controls (Table II).
Fig. 1.

The principle of SELEX experiments applied (a), and a histogram of LMNs calculated for the seven fragments selected multiple times by SELEX versus three sets of reference sequences (b). (a) To the surface of silica beads His-tag added FL4 bound through the metal nickel. Plasmids, pBluescript, either containing (P2 and P3) DNA fragments (white boxes) in the lacZ gene (blue edges separated) or not containing (P1), were bound by FL4, thereby selected. However, the site bound by FL4 can be positioned outside the cloning site (P1 and P3), i.e. contamination. The experiments were designed so that the number of type P2 was maximized. (b) A set of randomly combined 13 bps, another set where AAA is followed by randomly combined 9 bps, and a third set of 13 bps in the 5-3-5 arrangement. Ranks are labeled with the sum LMNs as well as the average, i.e. the sum divided by seven.

The principle of SELEX experiments applied (a), and a histogram of LMNs calculated for the seven fragments selected multiple times by SELEX versus three sets of reference sequences (b). (a) To the surface of silica beads His-tag added FL4 bound through the metal nickel. Plasmids, pBluescript, either containing (P2 and P3) DNA fragments (white boxes) in the lacZ gene (blue edges separated) or not containing (P1), were bound by FL4, thereby selected. However, the site bound by FL4 can be positioned outside the cloning site (P1 and P3), i.e. contamination. The experiments were designed so that the number of type P2 was maximized. (b) A set of randomly combined 13 bps, another set where AAA is followed by randomly combined 9 bps, and a third set of 13 bps in the 5-3-5 arrangement. Ranks are labeled with the sum LMNs as well as the average, i.e. the sum divided by seven.

The SELEX protocol optimized

When the concentration of MgSO4 was changed to 5, 25, 55, and 105 mM, respectively (Table II, Optimization I), in the presence of 150 mM KCl and 150 mM NaCl, the best W/B ratio, 1.70, was obtained with 105 mM MgSO4. However, the absolute number of white colonies, 17, was too small. Thus, the MgSO4 concentration of 55 mM was judged better: the primary W/B ratio was 1.09, and the secondary ratio relative to that found in the negative control was 1.63. When the KCl or NaCl concentration was increased to 350 mM in the presence of 5 mM MgSO4, the W/B ratio was not improved. In the presence of 55 mM MgSO4, when the NaCl concentration was increased stepwisely to 450, 550, 750, and 1,000 mM, respectively (Table II, Optimization 2), the best primary W/B ratio 1.93 was obtained with 1,000 mM NaCl: the secondary ratio was 3.39. When the experiment was repeated in the presence of 1,000 mM or 450 mM NaCl, and 55 mM MgSO4 at various temperatures, 60–95 °C (Table II, Optimization 3), the best primary W/B ratio, 0.71, and the best secondary ratio, 2.29, were obtained with 1,000 mM NaCl at 60 °C. On the basis of all these observations, the MgSO4 concentration of 50 mM, the KCl concentration of 150 mM, and the NaCl concentration of 1,000 mM were chosen for the final SELEX protocol at the temperature of 60 °C.

Selection of DNA fragments

In the first round of SELEX experiments (Table II), the primary and secondary W/B ratios observed were 1.19 and 2.28, respectively. In the second round the primary W/B ratio was lower, 1.34, and the secondary ratio was 1.34. Theoretically, the variation of plasmids selected in each round will decrease, thereby concentrating those containing binding-sites of FL4. After six more rounds, 89 clones were randomly chosen and sequenced. Among them were three copies of the same fragment: FL4-56 (the first entry in Table III). Two copies were found for six other fragments, FL4-2, FL4-25, FL4-26, FL4-29, FL4-40, FL4-74 (Table III, left top). Altogether these 15 copies formed 16.9% of the 89 clones sequenced. The other 74 fragments were of single copies: altogether 81 independent sequences were obtained.
Table III.

Fragments of DNA selected by SELEX

IDcopysubfrag.bpspositionsIDcopysubfrag.bpspositions
FL4-563671429253–1429319FL4-4211980444403–0444600
FL4-221710301019–0301189FL4-441I1010073230–0073330
FL4-252I590678436–0678494II1331164592–1164724
II2430206081–0206323FL4-451I940148653–0148746
FL4-262I2300639360–0639589II410206081–0206323
II560676358–0676413FL4-4611590346775–0346933
III371007989–1008025FL4-481490681987–0682035
FL4-292I1341040600–1040733FL4-4911700301019–0301188
II2400662925–0663164FL4-5013820648288–0648669
III860900779–0900864FL4-511911466318–1466408
FL4-4022101347230–1347439FL4-5211190524465–0524583
FL4-7423891604477–1604865FL4-531890793036–0793124
FL4-111421567873–1568014FL4-5411051540946–1541050
FL4-211700301019–0301188FL4-5512421515503–1515744
FL4-31420715037–0715078FL4-5811100075175–0075284
FL4-51I1071442314–1442420FL4-591311017813–1017843
II2420476169–0476410FL4-601I541287506–1287559
III731231109–1231181II663E. coli K12
FL4-612170292616–0292832FL4-6111301075739–1075868
FL4-811601165979–1166138FL4-6211101149647–1149756
FL4-91I340729947–0729980FL4-631860799661–0797746
II93no homologyFL4-641801287923–1288002
FL4-101I620214930–0214991FL4-6512741660946–1661219
II310022141–0022171FL4-661I2490709338–0709586
FL4-1112720048106–0048377II2020830728–0830929
FL4-121391617471–1617509FL4-681371027108–1027144
FL4-131651728886–1728940FL4-691941539656–1539749
FL4-1611141670785–1670898FL4-701521381522–1381573
FL4-1713711235147–1235517FL4-721I1260139160–0139285
FL4-1811020754580–0754681II1370019022–0019158
FL4-191I2440009691–0009934FL4-7512101347230–1347439
II1651042318–1042482FL4-7612141507767–1507980
FL4-2013041060276–1060579FL4-771I1531341065–1341217
FL4-2111090087812–0087920II2371731960–1732196
FL4-221800481286–0481365FL4-7911980052204–0052401
FL4-231500927463–0927512FL4-801640526230–0526293
FL4-241461588968–1589013FL4-811I1091507649–1507757
FL4-271I700888443–0888512II1151695509–1695623
II1111381434–1391544FL4-8214121506389–1506800
FL4-281I340792065–0792198FL4-8311180189560–0189677
II910637676–0637766FL4-841491467382–1467430
FL4-301451713156–1713200FL4-8611100348509–0348618
FL4-3111681140474–1140641FL4-8812191240256–1240474
FL4-321280375531–0375558FL4-9111520212718–0212869
FL4-331520072318–0072369FL4-9213580313462–0313819
FL4-341I1030026339–0026441FL4-9311141672618–1672731
II900481727–0481816FL4-941690322370–0322438
FL4-351341658671–1658704FL4-951I1870962153–0962339
FL4-3612111273014–1273224II1970142757–0142953
FL4-371530777830–0777882III1180657477–0659594
FL4-3811311471475–1471605IV760129288–0129363
FL4-411791681573–1681651FL4-9612700771434–0771703
Fragments of DNA selected by SELEX From another point of view, 14 of the 81 fragments were found as containing two independent subfragments each (subfragments I and II in Table III), 3 fragments (FL4-5, 26, 29) as containing three subfragments each, and yet another fragment (FL4-95) as containing four subfragments: altogether these forming 22.2% of the 81 fragments. With including fragments having single subfragments only, the average number of subfragments found in the 81 fragments was 1.28. Of all the 104 different subfragments, FL4-60-II was found originating in E. coli but not in P. OT3: a contamination. Another subfragment, FL4-9-I, was also a contamination but from an unknown origin. The average length of the subfragments was 185 bps. The average G: C content of the subfragments was 42%, which is the same as that of the genome of P. OT3. In what follows, the seven fragments selected multiple times, FL4-2, 25, 26, 29, 40, 56, 74, are further analyzed, since the possibility of their containing real binding sites is higher than that of other fragments selected only single time.

Discussion and analysis

Thirteen basepairs shared by the seven fragments with the minimum mismatches

Any DNA-binding domain (DBD) can cover only one side of DNA for ∼5 bps, and two such DBDs in a dimer are often separated by ∼10 bps or shorter along the DNA. Thus the sequence recognized by a dimer of a transcription factor will not much exceed ∼15 bps. Indeed, dimmers of FFRPs recognize 13 bps in the 5-3-5 arrangement (see Introduction). Thus, for each 13 bps randomly combined, the number of mismatches found with each of the seven fragments at its best resembling part was calculated: the least mismatch number, LMN (Fig. 2a).
Fig. 2.

Thirteen the best conserved among the seven DNA fragments selected multiple times by SELEX (a), and a selection of those in the 5-3-5 arrangements (b). The least mismatch numbers (LMNs) found with each fragment and their average are shown for each reference sequence.

Thirteen the best conserved among the seven DNA fragments selected multiple times by SELEX (a), and a selection of those in the 5-3-5 arrangements (b). The least mismatch numbers (LMNs) found with each fragment and their average are shown for each reference sequence. One of the three random 13 bps best conserved among the seven fragments selected multiple times was ATGAAAAAGTCAT, with the average LMN, 2.71 (Fig. 2a, highlighted in bold). This sequence is closely related with a 5-3-5 sequence, T, having only one mismatch: here bases the same as in G are underlined. In fact, ATGAA[AAA/TTT]TTCAT was found to be the single best 5-3-5 sequence conserved among the seven fragments with the average LMN of 3.14 (Fig. 2b, highlighted in bold). Another 5-3-5 sequence, which needs to be considered, is TTGAA[AAA/TTT]TTCAA. Many transcription factors auto-regulate the genes coding themselves, and so might be FL4. Upstream of the fl4 gene, the sequence TTGAAATTTACAA is positioned between a putative TATA box and an SD signal (Fig. 3). This is a typical formation for an FFRP to act as a repressor.[1),9)] The sequence resembles T more than ATT. The average LMN of 3.71 was calculated between TTGAATTTTTCAA and the seven selected fragments (Fig. 2b).
Fig. 3.

The nucleotide sequence of the region upstream of the fl4 gene in the P. OT3 genome. Candidates for a TATA box, an SD signal and the start codon are indicated. A putative FL4 binding site is also indicated with the number of mismatches with the sequence ATGAA[AAA/TTT]TTCAT. The arrow shows the direction of transcription.

The nucleotide sequence of the region upstream of the fl4 gene in the P. OT3 genome. Candidates for a TATA box, an SD signal and the start codon are indicated. A putative FL4 binding site is also indicated with the number of mismatches with the sequence ATGAA[AAA/TTT]TTCAT. The arrow shows the direction of transcription.

Possible repression of fl9 gene by FL4 protein

The average LMN found between ATGAA[AAA/TTT]TTCAT and the seven fragments was 3.14, which is better than our empirical threshold for theoretical identification of sites bound by FFRPs, ∼4. Yet FL4-25 and FL4-56 are on the border (Fig. 2b). Not all the seven fragments might contain sites functioning as real signal sequences. On the other hand, even when the score is below 4, the site might function as a signal sequence, when another binding site, even if it is less ideal, is positioned nearby: a cooperative interaction. When 7–8 bps are inserted between a pair of 13 bps, the two sites repeat with a periodicity of 20–21 bps, i.e. two full helical turns of DNA. With this arrangement, a pair of dimers can contact each other on the same side of the DNA, thereby forming a tetramer.[22)] More generally, the number of basepairs expected to be inserted is ∼[10-11] × N, where N is an integer. In FL4-56, and the part immediately downstream (Fig. 4f, shown by characters in lower case), three putative binding sites, two with four mismatches with ATGAA[AAA/TTT]TTCAT, and the other with five, were found as repeating with insertions of 18 and 20, respectively, positioned upstream of the gene pot1428536. The second and third sites were found as sandwiching a TATA-box, CTTAAAAA (Fig. 4f): a formation of repressing transcription of the gene. In FL4-2 (Fig. 4a), three putative binding sites, with three to five mismatches with ATGAA[AAA/TTT]TTCAT, were found repeating with insertions of 17 bps and 28 bps, respectively. The third site overlaps onto a TATA box, ATTGAATC, positioned upstream of gene pot0301583. This arrangement also represents a repression mode.
Fig. 4.

The nucleotide sequences of the seven fragments selected multiple times by SELEX experiments. Candidates for TATA boxes, SD signals and the start codons, ATG and TTG, are indicated. Putative FL4-binding sites are also indicated with the numbers of mismatches with the sequence ATGAA[AAA/TTT]TTCAT. Gene-coding regions are underlined. Except for (c), (e) and (g), possible modes of regulation, i.e. repression or activation, are written. In (c), (e) and (g) putative binding sites are found inside gene coding regions only, although still it is possible that they are designed for repression by FL4. In (b) two genes, pot0206233 and pot0206777, are overlapping onto each other inside an operon, and the 3′-end of the former and the 5′-end of the latter are indicated by parentheses,)) and (, respectively. Nucleotides found immediately outside the fragments are shown in lower case in (d), (f) and (g). For FL4-25 and 36 containing multiple subfragments, only single subfragments each are shown, since the other subfragments do not contain 13 bps closely related with ATGAA[AAA/TTT]TTCAT. For FL4-29, all the three sub-fragments are shown.

The nucleotide sequences of the seven fragments selected multiple times by SELEX experiments. Candidates for TATA boxes, SD signals and the start codons, ATG and TTG, are indicated. Putative FL4-binding sites are also indicated with the numbers of mismatches with the sequence ATGAA[AAA/TTT]TTCAT. Gene-coding regions are underlined. Except for (c), (e) and (g), possible modes of regulation, i.e. repression or activation, are written. In (c), (e) and (g) putative binding sites are found inside gene coding regions only, although still it is possible that they are designed for repression by FL4. In (b) two genes, pot0206233 and pot0206777, are overlapping onto each other inside an operon, and the 3′-end of the former and the 5′-end of the latter are indicated by parentheses,)) and (, respectively. Nucleotides found immediately outside the fragments are shown in lower case in (d), (f) and (g). For FL4-25 and 36 containing multiple subfragments, only single subfragments each are shown, since the other subfragments do not contain 13 bps closely related with ATGAA[AAA/TTT]TTCAT. For FL4-29, all the three sub-fragments are shown. Importantly, gene pot0301583 codes for another FFRP, FL9 (see Table I for the full ID). Immediately downstream of the fl9 gene, another gene codes DM2 in the same direction, most likely forming an operon. The protein DM2 is one of the three demi-FFRPs[20)] coded in the genome of P. OT3, having assembly domains only of full length FFRPs, e.g. FL4 and FL9. The two proteins, DM2 and FL9, are able to interact (Makino, K. et al., unpublished). These facts hint at the presence of a transcription network organized by FFRPs.

Possible transcription activation of pot1040906 by FL4.

Fragment FL4-29 was a chimera of three sub-fragments (Table III). It is not known which one of the three contains a binding site of FL4. In subfragment III the region upstream of gene pot1040906, coding a protein of an unknown function, was cloned (Fig. 4d), but the two other subfragments contain gene-coding regions only. In subfragment III two sites with five mismatches each with ATGAA[AAA/TTT]TTCAT were positioned with an insertion of 36 bps, which is not so different from 7 plus 10 multiplied by 3. Further upstream of the two sites, another site with three mismatches was positioned with an insertion of 9 bps, although not whole of this site was included in FL4-29III (shown in lower case in Fig. 4d). Downstream of the third site, separated by 3 bps, a putative TATA box is found. This particular arrangement fits well into the pattern of those predicted for activating transcription of genes by FFRPs,[1),5),9)] this time that of pot1040906. For the two other chimeric fragments, FL4-25 and 36, only single subfragments each are shown in Fig. 4, since the other subfragments do not contain 13 bps closely related with ATGAA[AAA/TTT]TTCAT.

Other statistical analyses

When LMNs were calculated between each of the 1,024 (i.e. 45) sequences in the 5-3-5 arrangement and the 73 fragments selected single time only (Fig. 5), the sequences, ATGAA[AAA/TTT]TTCAT and TTGAA[AAA/TTT]TTCAA, were given scores inside top 10%. The sum LMN was 347 and the average LMN was 4.75 with both 5-3-5 sequences. These observations are consistent with the idea that the 74 fragments are a mixture of those containing real binding sites, and a larger number of contaminants.
Fig. 5.

Thirteen bases the best conserved among 37 fragments selected single time only by SELEX. The sum and average least mismatch numbers (LMNs) found with the fragments are shown for each reference sequence.

Thirteen bases the best conserved among 37 fragments selected single time only by SELEX. The sum and average least mismatch numbers (LMNs) found with the fragments are shown for each reference sequence. The average of average LMNs calculated between a set of randomly combined 13 bps and the seven fragments was 5.00 (Fig. 1b). When reference sequences were restricted to those having AAA at the 5′ ends, or TTT at the 3′ ends, the secondary average of LMNs was improved to 4.79. This was because of the A:T content in the genome, 58%, which was higher than the average A:T of 50% in the first set and closer to that in the second set, 63%. The same A:T% was kept in a third set of 13 bps in the 5-3-5 arrangement. With this set, the average LMN was further improved to 4.51, suggesting that real-binding sites did have this type of arrangement. The scores calculated with the second set distributed more or less symmetrically to the two ends, but those calculated with the third set tailed more to the minimal mismatching end (Fig. 1b). The presence of a small number of 5-3-5 sequences having the smallest LMNs, i.e. ATGAA[AAA/TTT]TTCAT and its close variants, ATG[AAA/TTT]T and A[AAA/TTT]TTC, suggests that these sequences are indeed related with real binding sites of FL4.
  6 in total

1.  Crystallization and secondary-structure determination of a protein of the Lrp/AsnC family from a hyperthermophilic archaeon.

Authors:  N Kudo ; M D Allen ; H Koike ; Y Katsuya ; M Suzuki
Journal:  Acta Crystallogr D Biol Crystallogr       Date:  2001-03

2.  The archaeal feast/famine regulatory protein: potential roles of its assembly forms for regulating transcription.

Authors:  Hideaki Koike; Sanae A Ishijima; Lester Clowney; Masashi Suzuki
Journal:  Proc Natl Acad Sci U S A       Date:  2004-02-19       Impact factor: 11.205

3.  Systematic evolution of ligands by exponential enrichment: RNA ligands to bacteriophage T4 DNA polymerase.

Authors:  C Tuerk; L Gold
Journal:  Science       Date:  1990-08-03       Impact factor: 47.728

4.  Complete sequence and gene organization of the genome of a hyper-thermophilic archaebacterium, Pyrococcus horikoshii OT3.

Authors:  Y Kawarabayasi; M Sawada; H Horikawa; Y Haikawa; Y Hino; S Yamamoto; M Sekine; S Baba; H Kosugi; A Hosoyama; Y Nagai; M Sakai; K Ogura; R Otsuka; H Nakazawa; M Takamiya; Y Ohfuku; T Funahashi; T Tanaka; Y Kudoh; J Yamazaki; N Kushida; A Oguchi; K Aoki; H Kikuchi
Journal:  DNA Res       Date:  1998-04-30       Impact factor: 4.458

5.  Primer-directed enzymatic amplification of DNA with a thermostable DNA polymerase.

Authors:  R K Saiki; D H Gelfand; S Stoffel; S J Scharf; R Higuchi; G T Horn; K B Mullis; H A Erlich
Journal:  Science       Date:  1988-01-29       Impact factor: 47.728

6.  New metal chelate adsorbent selective for proteins and peptides containing neighbouring histidine residues.

Authors:  E Hochuli; H Döbeli; A Schacher
Journal:  J Chromatogr       Date:  1987-12-18
  6 in total
  1 in total

1.  The Lrp family of transcription regulators in archaea.

Authors:  Eveline Peeters; Daniel Charlier
Journal:  Archaea       Date:  2010-11-30       Impact factor: 3.273

  1 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.