| Literature DB >> 34282208 |
Peng-Fei Ke1,2,3, Dong-Sheng Xiong1,2,3, Jia-Hui Li1,2,3, Zhi-Lin Pan1,2,3, Jing Zhou1,2,3, Shi-Jia Li1,2,3, Jie Song1,2,3, Xiao-Yi Chen1,2,3, Gui-Xiang Li4,5, Jun Chen4,5, Xiao-Bo Li6, Yu-Ping Ning7,2, Feng-Chun Wu8,9, Kai Wu10,11,12,13,14,15,16,17.
Abstract
Finding effective and objective biomarkers to inform the diagnosis of schizophrenia is of great importance yet remains challenging. Relatively little work has been conducted on multi-biological data for the diagnosis of schizophrenia. In this cross-sectional study, we extracted multiple features from three types of biological data, including gut microbiota data, blood data, and electroencephalogram data. Then, an integrated framework of machine learning consisting of five classifiers, three feature selection algorithms, and four cross validation methods was used to discriminate patients with schizophrenia from healthy controls. Our results show that the support vector machine classifier without feature selection using the input features of multi-biological data achieved the best performance, with an accuracy of 91.7% and an AUC of 96.5% (p < 0.05). These results indicate that multi-biological data showed better discriminative capacity for patients with schizophrenia than single biological data. The top 5% discriminative features selected from the optimal model include the gut microbiota features (Lactobacillus, Haemophilus, and Prevotella), the blood features (superoxide dismutase level, monocyte-lymphocyte ratio, and neutrophil count), and the electroencephalogram features (nodal local efficiency, nodal efficiency, and nodal shortest path length in the temporal and frontal-parietal brain areas). The proposed integrated framework may be helpful for understanding the pathophysiology of schizophrenia and developing biomarkers for schizophrenia using multi-biological data.Entities:
Mesh:
Substances:
Year: 2021 PMID: 34282208 PMCID: PMC8290033 DOI: 10.1038/s41598-021-94007-9
Source DB: PubMed Journal: Sci Rep ISSN: 2045-2322 Impact factor: 4.379
Figure 1Flow chart of the brain network construction of the EEG signal. EEG electroencephalogram, PLV phase locking value. Figure (a) was generated by an EEG processing tool of “EEGLAB” (Version 2019.0, https://sccn.ucsd.edu/eeglab/index.php), based on MATLAB (Version R2018a). Figure (b–d) were generated by a brain network visualization tool of "BrainNet Viewer" (Version1.62, https://www.nitrc.org/projects/bnv/), based on MATLAB (Version R2018a).
Figure 2Overview of the proposed integrated machine learning framework for classifying schizophrenia. The proposed integrated machine learning framework for classifying schizophrenia consists of 5 M-methods. (a) Multi-biological data were collected from all subjects, including electroencephalogram (EEG) data, fecal data and blood data. (b) Multi-biological features were extracted from multi-biological data. (c) Multi-feature selection algorithms were used to eliminate redundant features, including recursive feature elimination (RFE), principal component analysis (PCA), and analysis of variance (ANOVA) (d) Multi-classifier were used to match heterogeneous biological features including support vector machine (SVM), random forest (RF), linear discriminant analysis (LDA), logistic regression (LR), and k-nearest neighbor (KNN) methods. (e) Multi-cross validation methods including tenfold, fivefold, threefold, and leave-one-out methods, were used to evaluate the performance of the trained model.
Figure 3Flowchart of the machine learning classification method.
Demographic and clinical characteristics used in the analysis.
| Characteristic | HCs (n = 50) | SZs (n = 49) | |
|---|---|---|---|
| Age, mean (SD) (years) | 41.7 (13.1) | 42.1 (12.5) | 0.89 |
| Male | 23 (46.0) | 24 (49.0) | 0.77 |
| Female | 27 (54.0) | 25 (51.0) | |
| Education years, mean (SD) (years) | 14.2 (3.6) | 11.6 (3.4) | < 0.001 |
| PANSS, mean (SD) | NA | 58.84 (17.50) | NA |
HCs healthy controls, SZs patients with schizophrenia, NA not applicable.
Classification performance of the optimal model including different input features using the integrated machine learning framework (tenfold).
| Input feature | Feature Selection Method | Classifier | Accuracy (%) | Sensitivity (%) | Specificity (%) | AUC | |
|---|---|---|---|---|---|---|---|
| Gut microbiota features (n = 77) | RFE | RF | 70.8 | 58.3 | 83.3 | 0.80 | 0.03 |
| Blood features (n = 12) | Noneb | KNN | 83.3 | 83.3 | 83.3 | 0.88 | 0.010 |
| EEG features (n = 574) | RFE | RF | 79.2 | 83.3 | 75.0 | 0.90 | 0.010 |
| Combined features (n = 663) | None | SVM | 91.7 | 91.7 | 91.7 | 0.97 | 0.010 |
AUC area under the receiver operating characteristic curve, RFE recursive feature elimination, KNN k-nearest neighbor, LR logistic regression, RF random forest, SVM support vector machine, EEG electroencephalogram.
aThe statistical significance of the permutation test was set to p < 0.05.
bNone means no feature selection algorithm was used.
Figure 4Areas under the receiver operating characteristic curves (AUC) for the best model comparing the gut microbiota features, blood features, electroencephalogram features and the combination of GMV, BF and EF as the input for machine learning. Each curve in the figure represents the ROC curve of the best model using different input features. GMF gut microbiota features, BF blood features, EF electroencephalogram features, CF combined features. This figure was generated by “Visual Studio Code” (Version 1.56, https://code.visualstudio.com/).
Top 34 features (5%) showing the most discriminative biomarkers for multi-biological predictions.
| Number | Feature name | Feature type | Number | Feature name | Feature type |
|---|---|---|---|---|---|
| 1 | SOD | Blood | 18 | GM | |
| 2 | MLR | Blood | 19 | GM | |
| 3 | GM | 20 | PLT | Blood | |
| 4 | MON | Blood | 21 | alpha2_aNLe_P4 | EEG |
| 5 | GM | 22 | GM | ||
| 6 | GM | 23 | beta1_aLambda | EEG | |
| 7 | NEU | Blood | 24 | GM | |
| 8 | CRP | Blood | 25 | GM | |
| 9 | GM | 26 | GM | ||
| 10 | theta_aNLe_T6a | EEG | 27 | GM | |
| 11 | theta_aNe_T6 | EEG | 28 | theta_aDc_FP1 | EEG |
| 12 | theta__aNCp_T6 | EEG | 29 | alpha2__aNCp_P4 | EEG |
| 13 | WBC | Blood | 30 | beta2_aNLe_FP2 | EEG |
| 14 | NLR | Blood | 31 | beta2_aDc_O2 | EEG |
| 15 | GM | 32 | GM | ||
| 16 | gamma_aDc_F7 | EEG | 33 | alpha2_aNLe_T4 | EEG |
| 17 | GM | 34 | alpha2__aNCp_T4 | EEG |
The top 34 features are listed in the descending order of their weights.
GM gut microbiota, EEG electroencephalogram, SOD superoxide dismutase, MLR monocyte–lymphocyte ratio, MON monocyte, NEU neutrophil, CRP C-reactive protein, WBC white blood cell, NLR neutrophil–lymphocyte ratio, PLT platelet, aNLe nodal local efficiency, aNe nodal efficiency, aNCp nodal clustering coefficient, aDc degree centrality.
aThe EEG features are represented as a_b_c, where a represents the frequency band, b represents brain network attributes, and c represents the electrode channel.
bUndefined Lachnospiraceae.
cUndefined Ruminococcaceae.
Comparison of classification performance with existing research.
| References | Sample size | Input feature | Feature selection method | Classifier | Cross validation method | Performance |
|---|---|---|---|---|---|---|
| Shen et al.[ | SZ = 64 HC = 53 | Gut microbiota | Boruta variable selection | RF | None | AUC = 0.837 |
| Brisa et al.[ | SZ = 58 HC = 123 | Blood and cognitive | PLS-DA | LDA | Tenfold | Accuracy = 0.86 AUC = 0.89 |
| Jason et al.[ | SZ = 40 HC = 12 | EEG | None | SVM | None | Accuracy = 0.87 Sensitivity = 0.90 Specificity = 0.77 |
| Sai Krishna Tikka et al.[ | SZ = 38 HC = 20 | EEG | None | SVM | Hold-out | Accuracy = 0.79 Sensitivity = 0.92 Specificity = 0.50 AUC = 0.71 |
| Our best | SZ = 49 HC = 50 | Gut microbiota Blood EEG | None | SVM | Tenfold | Accuracy = 0.92 Sensitivity = 0.92 Specificity = 0.92 AUC = 0.97 |
RF random forest, PLS-DA partial least squares discriminant analysis, LDA linear discriminant analysis, EEG electroencephalogram, SVM support vector machine, AUC area under the receiver operating characteristic curve.