| Literature DB >> 34198491 |
Yaxuan Liu1, Olga Axell1, Tom van Leeuwen2, Robert Konrat3, Pedram Kharaziha1, Catharina Larsson1, Anthony P H Wright2, Svetlana Bajalica-Lagercrantz1.
Abstract
Rare germline pathogenic TP53 missense variants often predispose to a wide spectrum of tumors characterized by Li-Fraumeni syndrome (LFS) but a subset of variants is also seen in families with exclusively hereditary breast cancer (HBC) outcomes. We have developed a logistic regression model with the aim of predicting LFS and HBC outcomes, based on the predicted effects of individual TP53 variants on aspects of protein conformation. A total of 48 missense variants either unique for LFS (n = 24) or exclusively reported in HBC (n = 24) were included. LFS-variants were over-represented in residues tending to be buried in the core of the tertiary structure of TP53 (p = 0.0014). The favored logistic regression model describes disease outcome in terms of explanatory variables related to the surface or buried status of residues as well as their propensity to contribute to protein compactness or protein-protein interactions. Reduced, internally validated models discriminated well between LFS and HBC (C-statistic = 0.78-0.84; equivalent to the area under the ROC (receiver operating characteristic) curve), had a low risk for over-fitting and were well calibrated in relation to the known outcome risk. In conclusion, this study presents a phenotypic prediction model of LFS and HBC risk for germline TP53 missense variants, in an attempt to provide a complementary tool for future decision making and clinical handling.Entities:
Keywords: Li-Fraumeni syndrome; germline TP53 missense variants; hereditary breast cancer; protein conformation; quantitative prediction model
Mesh:
Substances:
Year: 2021 PMID: 34198491 PMCID: PMC8231809 DOI: 10.3390/ijms22126345
Source DB: PubMed Journal: Int J Mol Sci ISSN: 1422-0067 Impact factor: 5.923
Explanatory variables related to protein conformation.
| Protein Characteristics | Variables | Predictor Algorithm [Ref] |
|---|---|---|
|
| ||
| buried/surface | Bur | Pymol-findSurfaceResidues script [ |
|
| ||
| disorder (trained on Disprot DB) | disprot_dif | Espritz [ |
| disorder (trained on NMR structures) | nmr_dif | Espritz [ |
| disorder (trained on X-ray structures) | xray_dif | Espritz [ |
| disorder (longer regions) | iupl_dif | IUPred2A [ |
| disorder (short regions) | iups_dif | IUPred2A [ |
|
| ||
| protein backbone dynamics | dyn_dif | Dynamine [ |
|
| ||
| alpha-helix/beta-sheet | Sec_dif | Meta-structure [ |
|
| ||
| protein compactness | comp_dif | Meta-structure [ |
| protein globularity | iupstr_dif | IUPred2A [ |
|
| ||
| protein protein interaction | PPI6_dif | Meta-structure-PPI [ |
| protein protein interaction | anc_dif | IUPred2A [ |
Figure 1Location of TP53 missense variants in the TP53 protein sequence and in relation to its predicted disorder. (a). Schematic illustration of the TP53 amino acid sequence and protein domains with the location of the 24 LFS variants shown above (cyan green-blue rhombus) and the 24 HBC-variants indicated below (magenta purple-red circles) [39]. The TP53 domains are illustrated for the transactivation domain (TAD), the proline-rich region (PRR), the DNA binding domain (DBD), the nuclear localization signal (NLS), the oligomerization domain (OD) and the C-terminal regulatory domain (CTD). (b) Predicted disorder profile of wild type TP53. The IUPred2A predictor was used with the “long” argument. Scores > 0.5 (above dotted line) indicate disordered regions. The approximate location of the DBD and OD are shown (grey shading). (c) Predicted compactness of wild type TP53. The dotted line at a value of 250 (y-axis) emphasizes the higher compactness values predicted for the DBD. The approximate location of the DBD and OD are shown (grey shading). (d) Predicted secondary structure of wild type TP53. Values > 0 (dotted line) are predicted to be alpha-helical and values < 0 are predicted to have beta-strand conformation. The approximate location of the DBD and OD are shown (grey shading). (e) Predicted regions with protein interaction propensity in wild type TP53. The dotted line shows a level equivalent to 5% of the maximum value. Apparently artefactual values for the first 4 residues and last 3 residues of TP53 were omitted. The approximate location of the DBD and OD are shown (grey shading).
Figure 2Comparison of surface exposure of TP53 missense variants in LFS and HBC. (a) LFS-residues are significantly enriched in Buried residues with lower surface exposure (
Figure 3Protein conformation parameters are associated with disease phenotype and may have predictive value. (a) Multivariate logistic regression models for prediction of phenotype class (LFS or HBC) using a range of available protein conformation related explanatory variables describing different protein conformation aspects (full model, mod_full). A reduced model (mod_4v) was produced by stepwise variable exclusion from the full model (rms package). Further reduction was done by progressive manual removal of the least well performing variable to produce models with 3 and 2 explanatory variables, respectively (mod_3v and mod_2v). Indicators of model performance (C-statistic) are shown. (b) Bimodal probability distributions for models, showing the overall separation of output variables (LFS and HBC). Histograms show the distribution of the predicted probabilities of residues causing LFS, for the different models. The overall separation of variants as LFS (cyan line) or HBC (magenta line) are shown for each model. (c) Probability values for individual LFS (cyan) and HBC (magenta) variants produced by the full model and reduced models (columns 1–4 in each panel) as well as leave-one-out cross validation results (loo_cv) in which each respective variant is left out from a reduced model that is then used to predict the outcome associated with the left-out variable (column 5 in each panel). An asterisk (*) indicates that the loo_cv model contains the same variables as the mod_4v model (model details and results for each of the loo_cv model are tabulated in Table S5). Based on the binomial distribution minima in part b, probability values >0.5 for LFS and <0.5 for HBC are colored darker to give an indication of the relative performance (correct predictions) of the different models as well as how performance is affected in the cross-validation procedure in which predictions are made for each individual variant by models excluding data for the predicted variant. Variant/model combinations with lighter color indicate incorrect predictions.
Reduced models for multivariate logistic regression analysis of disease outcome.
| Intercept and Variable | β | OR (95% CI) | |
|---|---|---|---|
|
| |||
| Intercept | 1.282 |
| |
| Bur | |||
| Surface vs. Buried | −2.465 | 0.09 (0.02–0.44) |
|
| unknown vs. Buried | −9.897 | 5.03 × 10−5(4.31 × 10−27–5.87× 1017) | 0.703 |
| comp_dif | 0.023 | 3.88 (1.30–11.59) |
|
| PPI6_dif | 0.001 | 1.29 (0.99–1.67) | 0.059 |
| nmr_dif | 10.461 | 2.20 (0.79–6.11) | 0.132 |
|
| |||
| Intercept | 1.094 |
| |
| Bur | |||
| Surface vs. Buried | −2.229 | 0.11 (0.02–0.50) |
|
| unknown vs. Buried | −9.03 | 0.00012 (1.18 × 10−26–1.22 × 1018) | 0.727 |
| comp_dif | 0.015 | 2.40 (1.08–5.30) |
|
| PPI6_dif | 0.0006 | 1.15 (0.95–1.39) | 0.165 |
|
| |||
| Intercept | 0.914 |
| |
| Bur | |||
| Surface vs. Buried | −1.886 | 0.15 (0.04–0.62) |
|
| unknown vs. Buried | −8.804 | 0.0002 (1.48 × 10−26–1.53 × 1018) | 0.733 |
| comp_dif | 0.014 | 2.30 (1.04–5.09) |
|
β = beta coefficient; OR = odds ratio; CI = confidence interval; Bur = categorical variable describing whether residues are burried, surface or of unknown location in the TP53 tertiary structure; comp_dif = a continuous variable showing the effect of each variant on the predicted compactness of TP53 at the location of each variant residue (variant value minus wild type value); PPI6_dif and nmr_dif are continuous variables calculated as for comp_dif but reflecting the effect of variant residues on the predicted protein interaction protensity and the predicted intrinsic disorder of TP53, respectively; * = p < 0.05.
Figure 4Potential for the most favored model. (a) Decision curve analysis of models for prediction of phenotypic outcome (LFS or HBC). The y-axis indicates the net benefit of using mod_3v model (red line). The thin gray line (All) shows net benefit values expected assuming all assessed variants are LFS. The darker gray line (None) shows net benefit values expected assuming no assessed variants are LFS. The net benefit for prediction of LFS variants is regarded as positive for probability values exceeding those for the “All” and “None” values. (b) The nomogram that facilitates manual estimation of the risk of LFS disease outcome using the mod_3v model. For the value of each explanatory variable the equivalent value on the “Points” scale is assessed. The sum of all Points values is then located on the “Total Points” scale (middle green row) so that the corresponding probability value can be read from the “LFS outcome rate” scale (upper green row). The dotted arrows show a hypothetical example for a surface residue with a comp_dif value of 50 and a PPI6_dif value of −500, for which the equivalent “Point” values (10, 40 and 75) summate to 125 (“Total Points”), giving a LFS outcome risk of slightly over 0.3.