| Literature DB >> 36230750 |
Byung-Hoon Kim1,2, Hyeonhoon Lee3,4,5, Kyu Sung Choi5,6, Ju Gang Nam5,6, Chul-Kee Park7, Sung-Hye Park8, Jin Wook Chung5,6,9, Seung Hong Choi5,6,10.
Abstract
O6-methylguanine-DNA methyl transferase (MGMT) methylation prediction models were developed using only small datasets without proper external validation and achieved good diagnostic performance, which seems to indicate a promising future for radiogenomics. However, the diagnostic performance was not reproducible for numerous research teams when using a larger dataset in the RSNA-MICCAI Brain Tumor Radiogenomic Classification 2021 challenge. To our knowledge, there has been no study regarding the external validation of MGMT prediction models using large-scale multicenter datasets. We tested recent CNN architectures via extensive experiments to investigate whether MGMT methylation in gliomas can be predicted using MR images. Specifically, prediction models were developed and validated with different training datasets: (1) the merged (SNUH + BraTS) (n = 985); (2) SNUH (n = 400); and (3) BraTS datasets (n = 585). A total of 420 training and validation experiments were performed on combinations of datasets, convolutional neural network (CNN) architectures, MRI sequences, and random seed numbers. The first-place solution of the RSNA-MICCAI radiogenomic challenge was also validated using the external test set (SNUH). For model evaluation, the area under the receiver operating characteristic curve (AUROC), accuracy, precision, and recall were obtained. With unexpected negative results, 80.2% (337/420) and 60.0% (252/420) of the 420 developed models showed no significant difference with a chance level of 50% in terms of test accuracy and test AUROC, respectively. The test AUROC and accuracy of the first-place solution of the BraTS 2021 challenge were 56.2% and 54.8%, respectively, as validated on the SNUH dataset. In conclusion, MGMT methylation status of gliomas may not be predictable with preoperative MR images even using deep learning.Entities:
Keywords: MRI; O6-methylguanine-DNA methyl transferase; gliomas; neural network; radiogenomics
Year: 2022 PMID: 36230750 PMCID: PMC9562637 DOI: 10.3390/cancers14194827
Source DB: PubMed Journal: Cancers (Basel) ISSN: 2072-6694 Impact factor: 6.575
Figure 1Patient inclusion and exclusion criteria. Abbreviations: T1w, T1-weighted imaging; T2w, T2-weighted imaging; T1wCE, contrast-enhanced T1-weighted imaging; FLAIR, fluid-attenuated inversion recovery.
Patient demographics and genetic information.
| Number of Patients | Age, Mean ± SD (Years) | PFS, Median (95% CI) (Days) | ||
|---|---|---|---|---|
| Sex | ||||
| Male | 240 (60%) | 52.6 ± 15.7 | 327 (287–372) | 0.655 |
| Female | 160 (40%) | 51.9 ± 14.7 | 362 (301–481) | |
| MGMT | ||||
| Unmethylated | 203 (50.8%) | 52.6 ± 15.6 | 396 (328–526) | <0.0001 * |
| Methylated | 197 (49.2%) | 52.0 ± 14.9 | 974 (698–1302) |
Abbreviations: SD = standard deviation; CI = confidence interval; MGMT = O6-methylguanine-DNA methyltransferase; PFS= progression-free survival. * indicates the p value for a significant difference in PFS using the log-rank test.
Comparison of previous prediction models of MGMT methylation.
| Previous Study | Dataset | MR Sequence | Input Feature | Model Architecture | Dimension | Diagnostic Performance |
|---|---|---|---|---|---|---|
| Han et al. [ | TCIA ( | T1w, T2w, FLAIR | Raw images | CRNN | 2D axial CNN with RNN in slice-direction (z-axis) | Acc 67% (validation), 62% (test), precision (67%), recall (67%) |
| Sasaki et al. [ | Osaka International Cancer Institute ( | T1w, T2w, FLAIR, T1wCE | Radiomics | Supervised principal component analysis | 3D VOI of 1 mm isotropic resampled image | Acc 67% (mean by 10-fold cross-validation) |
| Levner et al. [ | Tom Baker Cancer Centre ( | T2w, FLAIR, | Texture analysis | L1-regularized neural network | 2D axial | Acc 87.7% |
| Drabycz et al. [ | Tom Baker Cancer Centre ( | T2w, FLAIR, | Texture analysis | Linear discriminant analysis | 2D axial | Acc 71% |
| Yogananda et al. [ | TCIA ( | T2w | Raw images | 3D-DenseUNet | 3D patch | Acc 94.7% (mean by 3-fold cross-validation) |
| Wei et al. [ | Shanxi Medical University ( | T1wCE, FLAIR, ADC | Radiomics | Logistic regression | 3D VOI | Acc 77% (validation; |
| Korfiatis et al. [ | Mayo Clinic ( | T2w, T1wCE | Texture analysis | Support vector machines, random forest classifiers | 2D ROI | AUC 0.85 |
Comparison of model performance using different models and sequences in validation sets.
| Dataset | CNN Architecture | MR Sequence | Metrics † | |||
|---|---|---|---|---|---|---|
| Best AUROC (%) | Accuracy (%) | Precision (%) | Recall (%) | |||
| Experiment 1 | EfficientNet-B0 | FLAIR-T1wCE-T2w-T1w | 46.4 | 55.9 | 55.9 | 83.2 |
| Experiment 2 | SEResNeXt50 | FLAIR-T1wCE | 57.8 | 55.5 | 36.6 | 44.0 |
| Experiment 3 | SEResNet50 | T2w | 54.9 | 57.1 | 61.6 | 68.7 |
† Metrics were calculated using the validation set for each experiment. The mean and standard deviation were obtained from five different models of the same CNN architectures and MR sequences trained using five different seed numbers. Data are given as the mean standard deviation (range).
Figure 2Model performance in the (a) validation (or tuning) and (b) test sets. For each of the experiments, both accuracy and AUROC are shown to report the model performance in the (a) validation (i.e., tuning) and (b) test sets. Note that the dashed red lines are the chance level (50%). The horizontal axis is the dataset on which the model was trained/validated (i.e., trained/tuned): “Public” indicates that the model was trained/validated on the BraTS dataset and tested on the SNUH dataset. “SNUH” indicates that the model was trained/validated on the SNUH dataset and tested on the BraTS dataset. “Merged” indicates that the model was trained/validated and tested with a randomly split SNUH + BraTS dataset. Error bars indicate the standard deviation of the metrics. Note that the validation metrics are better than the test metrics because the model training was stopped early according to the high validation accuracy. The red dotted lines indicate the chance level. Abbreviations: AUROC, area under the receiver operating characteristic curve.
Comparison of model performance using different models and sequences in test sets.
| Dataset | CNN Architecture | MR Sequence | Metrics † | |||
|---|---|---|---|---|---|---|
| Best AUROC (%) | Accuracy (%) | Precision (%) | Recall (%) | |||
| Experiment 1 | EfficientNet-B0 | FLAIR-T1wCE-T2w-T1w | 51.6 | 49.8 | 49.0 | 80.4 |
| Experiment 2 | SEResNeXt50 | FLAIR-T1wCE | 51.7 | 51.9 | 54.6 | 62.0 |
| Experiment 3 | SEResNet50 | T2w | 51.5 | 50.4 | 32.6 | 38.2 |
| Experiment 4 | 3D-ResNet | T1wCE | 56.2 | 54.8 | 53.6 | 59.9 |
† Metrics were calculated using the test set for each experiment. The mean and standard deviation were obtained from five different models of the same CNN architectures and MR sequences trained using five different seed numbers. Data are given as the mean standard deviation (range).
Figure 3Probability score distribution according to the MGMT labels. The probability scores were predicted from the best model of each experiment (specified in Table 4) and obtained using the test set specific to each experiment. Note that there is no noticeable boundary in the distribution of data points between the groups with high and low probability scores, according to MGMT labels, which are indicated as different colors. MGMT+ (blue) indicates the methylated MGMT promotor group, and MGMT- (orange) indicates the unmethylated MGMT promotor group. The red dotted lines indicate chance level. Abbreviations: MGMT, O6-methylguanine-DNA methyltransferase.