| Literature DB >> 27193998 |
Miroslav Krepl1, Antoine Cléry2, Markus Blatter3, Frederic H T Allain4, Jiri Sponer5.
Abstract
RNA recognition motif (RRM) proteins represent an abundant class of proteins playing key roles in RNA biology. We present a joint atomistic molecular dynamics (MD) and experimental study of two RRM-containing proteins bound with their single-stranded target RNAs, namely the Fox-1 and SRSF1 complexes. The simulations are used in conjunction with NMR spectroscopy to interpret and expand the available structural data. We accumulate more than 50 μs of simulations and show that the MD method is robust enough to reliably describe the structural dynamics of the RRM-RNA complexes. The simulations predict unanticipated specific participation of Arg142 at the protein-RNA interface of the SRFS1 complex, which is subsequently confirmed by NMR and ITC measurements. Several segments of the protein-RNA interface may involve competition between dynamical local substates rather than firmly formed interactions, which is indirectly consistent with the primary NMR data. We demonstrate that the simulations can be used to interpret the NMR atomistic models and can provide qualified predictions. Finally, we propose a protocol for 'MD-adapted structure ensemble' as a way to integrate the simulation predictions and expand upon the deposited NMR structures. Unbiased μs-scale atomistic MD could become a technique routinely complementing the NMR measurements of protein-RNA complexes.Entities:
Mesh:
Substances:
Year: 2016 PMID: 27193998 PMCID: PMC5291263 DOI: 10.1093/nar/gkw438
Source DB: PubMed Journal: Nucleic Acids Res ISSN: 0305-1048 Impact factor: 16.971
Figure 1.The studied RRM protein/RNA complexes: (A) Fox-1 complex. The non-canonical (hydrophobic pocket) and canonical parts of the protein/RNA interface are highlighted in red and blue respectively; (B) SRSF1 complex. The protein/RNA interface is highlighted in red. Other parts of the RNA molecule do not form specific interactions with the protein; the secondary structure of the proteins is labeled and highlighted in purple (α-helices), yellow (β-sheets) and cyan/white (loops). The RNA backbone is traced in brown. The nucleotides are numbered and the chain termini labeled. For additional structural details see Supplementary Figures S1 and S2.
List of simulations
| Simulation namea,b | NMR restraints initially applied | Length (ns) |
|---|---|---|
| 2err_14_1 | No | 1000 |
| 2err_14_2 | No | 1000 |
| 2err_14_rst1 | Yes | 1000 |
| 2err_14_rst2 | Yes | 1000 |
| 2err_14_rst3 | Yes | 1000 |
| 2err_12_1 | No | 1000 |
| 2err_12_2 | No | 1000 |
| 2err_12_rst1 | Yes | 1000 |
| 2err_12_rst2 | Yes | 1000 |
| 2err_12_rst3 | Yes | 1000 |
| 2err_12_rst4 | Yes | 1000 |
| 2err_99_1 | No | 1000 |
| 2err_99_2 | No | 1000 |
| 2err_99_rst1 | Yes | 1000 |
| 2err_99_rst2 | Yes | 1000 |
| 2err_99_rst3 | Yes | 1000 |
| 2err_99_rst4 | Yes | 1000 |
| 2m8d_14_1 | No | 1000 |
| 2m8d_14_2 | No | 1000 |
| 2m8d_14_rst1 | Yes | 1000 |
| 2m8d_14_rst2 | Yes | 1000 |
| 2m8d_14_rst3 | Yes | 1000 |
| 2m8d_14_rst4 | Yes | 1000 |
| 2m8d_12_1 | No | 1000 |
| 2m8d_12_2 | No | 1000 |
| 2m8d_12_rst1 | Yes | 1000 |
| 2m8d_12_rst2 | Yes | 1000 |
| 2m8d_12_rst3 | Yes | 1000 |
| 2m8d_12_rst4 | Yes | 1000 |
| 2m8d_12_rst5 | Yes | 1000 |
| 2m8d_99_1 | No | 700 |
| 2m8d_99_2 | No | 600 |
| 2m8d_99_rst1 | Yes | 1000 |
| 2m8d_99_rst2 | Yes | 1000 |
| 2m8d_99_rst3 | Yes | 1000 |
| 2m8d_99_rst4 | Yes | 1000 |
| 2m8d_99_rst5 | Yes | 1000 |
| 2m8d_14_short1c | Yes | 1000 |
| 2m8d_14_short2c | Yes | 1000 |
| 2m8d_12_short1c | Yes | 1000 |
| 2m8d_12_short2c | Yes | 1000 |
| 2m8d_99_short1c | Yes | 2000 |
| 2m8d_99_short2c | Yes | 1000 |
| 2m8d_14_R142Ad | No | 1000 |
| 2m8d_12_R142Ad | No | 4000 |
| 2m8d_12_R142A_2d | No | 2000 |
| 2m8d_12_R142A_TI_1e | No | 54 × 50 |
| 2m8d_12_R142A_TI_2e | No | 54 × 200 |
aAfter 120 ns of the simulation (unrestrained part, see Materials and Methods section), all of the initially restrained trajectories (marked as ‘_rst’) are fully independent simulation runs. However, up to 120 ns, some of them share a common restrained part of the trajectory. Full explanation is in the Supplementary Scheme S1.
bThe ‘14’, ‘12’ and ‘99’ numerals in the simulation name indicate ff14SB, ff12SB and ff99SB protein force field versions, respectively. For the RNA, the ff99bsc0χOL3 force field was used in all simulations.
cThe nucleotides 1–3 and amino-acids 106–114 were removed.
dThe R142A mutation was introduced into the system by molecular modeling, with the final structure of the 2m8d_12_rst1 simulation used as the starting configuration.
eBoth TI calculations consist of 54 independent simulations, each lasting either 50 (first simulation run) or 200 ns (second simulation run).
Number of protein–RNA NOE distances that are satisfied in the simulations of the Fox-1 complex for the individual nucleotides (the number of the observed NOEs is given on the first line). Averaged values of weighted NOE distances (see Materials and Methods) calculated over the entire simulation trajectories and over the last 50 ns are used. Please see Supplementary Figures S5 and S6 for structure visualization of the NOE violations
|
|
Number of protein–RNA NOE distances that are satisfied in the simulations of the SRSF1 complex for the individual nucleotides (the number of the observed NOEs is given on the first line). Averaged values of weighted NOE distances (see Materials and Methods) calculated over the entire simulation trajectories and over the last 50 ns are used. Please see Supplementary Figure S7 for structure visualization of the NOE violations
|
|
Figure 2.Time development of heavy atom distances of specific intermolecular H-bond interactions in selected Fox-1 protein / RNA complex (PDB: 2err) simulations: 1. U1(N3)/Ser155(O); 2. U1(O2)/Asn151(ND2); 3. G2(N1)/Ile124(O); 4. G2(N2)/Ile124(O); 5. C3(N3)/Asn151(ND2); 6. C3(N4)/Ser155(O); 7. U5(N3)/Asn190(O); 8. U5(O2)/Thr192(N); 9. G6(N1)/Thr192(O); 10.G6(O6)/Arg118(sc); 11. G6(N7)/Arg118(sc). The H-bonds written in bold are present in the experimental NMR ensemble structures with over 90% occurrence. The H-bonds ‘1’ and ‘2’ are absent in the NMR ensemble but are often formed during the simulations. The ‘sc’ abbreviation for Arginine indicates any of the three side-chain nitrogen atoms potentially acting as donors in an H-bond. The bond angles were monitored to verify that short interatomic distances correspond to H-bonding but are not shown for space reasons. The remaining simulations are summarized in Supplementary Figure S4.
Figure 3.Time development of heavy atom distances of specific intermolecular H-bond interactions in selected SRSF1 protein / RNA complex (PDB: 2m8d) simulations: 1. G5(N1)/Ala150(O); 2. G5(O6)/Ala150(N); 3. G6(N1)/Asp139(OD); 4. G6(N2)/Asp139(OD); 5. G6(O4')/Gln135(NE2); 6. G6(O6)/Arg142(sc); 7. A7(N6)/Asp136(OD); 8. A7(N1)/Ser133(OG). The H-bonds written in bold are present in the experimental NMR ensemble structures with over 90% occurrence. The H-bond ‘6’ is present in just one structure of the NMR ensemble but it is frequently observed in our simulations. The remaining simulations are summarized in Supplementary Figure S3.
Figure 4.The initial arrangement (the first NMR frame, top left) of the U1/G2/C3/Phe126 hydrophobic pocket and the three alternative conformations seen in the simulations during time periods where the C3 nucleotide was stably bound. The Sim1 conformation was the most common while the others were less frequent. The H-bonds are indicated by dotted lines between heavy atoms. The Table summarizes the stacking interactions, H-bonds, and the number of satisfied protein–RNA intermolecular NOE distances in the specific conformations. PDB files of representative structures can be found in Supporting Information.
Figure 5.Fox-1 complex. The Arg118 side chain is forming H-bonds with the G6 base as either Arg118(NE)/G6(O6) (left), the Arg118(NH1)/G6(N7) (middle), or the Arg118(NH2)/G6(N7) and Arg118(NE)/G6(O6) interactions (right). The first arrangement is populated only in restrained parts of the simulations while the other two are populated in the unrestrained parts.
Figure 6.(A) Overlap of G5 and Trp134 aromatic rings in the NMR (top) and in the simulations (bottom). This change, while minor, usually resulted into at least one G5/Trp134 NOE distance violation greater than 1 Å. (B) In simulations of the SRSF1 complex, the Lys138 side chain fluctuated between G5 (top) and G6 (bottom) Hoogsteen base edges. The typical heavy atom distances are shown (in Å). (C) The Arg142 side chain was often simultaneously interacting with G6 and Asp139 residues in the SRSF1 simulations, effectively increasing the protein's specificity for the guanine in this position by simultaneously recognizing the entire Watson-Crick edge of the base in a highly specific way.
Figure 7.(A) NMR titrations of the 15N-labeled GB1-SRSF1 RRM2 WT and R142A proteins with the unlabeled 5′-AGGAC-3′ RNA. The peaks corresponding to the free WT protein are colored blue. The 1:1 RNA-bound proteins (with WT or mutant protein) are colored green and red, respectively. The differences in chemical shift perturbations observed upon RNA binding are indicated by black arrows. (B) ITC data recorded with SRSF1 RRM2 WT and R142A proteins in the presence of the 5′-AGGAC-3′ RNA. The estimated Kd values are shown.