| Literature DB >> 35271184 |
Mohammad S Jassas1, Qusay H Mahmoud1.
Abstract
Modern applications, such as smart cities, home automation, and eHealth, demand a new approach to improve cloud application dependability and availability. Due to the enormous scope and diversity of the cloud environment, most cloud services, including hardware and software, have encountered failures. In this study, we first analyze and characterize the behaviour of failed and completed jobs using publicly accessible traces. We have designed and developed a failure prediction model to determine failed jobs before they occur. The proposed model aims to enhance resource consumption and cloud application efficiency. Based on three publicly available traces: the Google cluster, Mustang, and Trinity, we evaluate the proposed model. In addition, the traces were also subjected to various machine learning models to find the most accurate one. Our results indicate a significant correlation between unsuccessful tasks and requested resources. The evaluation results also revealed that our model has high precision, recall, and F1-score. Several solutions, such as predicting job failure, developing scheduling algorithms, changing priority policies, or limiting re-submission of tasks, can improve the reliability and availability of cloud services.Entities:
Keywords: Google cluster trace; Mustang trace; Random Forest (RF); Trinity trace; cloud computing; failure prediction; fault tolerance
Mesh:
Year: 2022 PMID: 35271184 PMCID: PMC8914926 DOI: 10.3390/s22052035
Source DB: PubMed Journal: Sensors (Basel) ISSN: 1424-8220 Impact factor: 3.576
The state of the art in the field of cloud computing for failure analysis and prediction.
| References | Trace | Focus | Model | Results |
|---|---|---|---|---|
| Chen Xin et al. [ | Google trace | Predicting job failure | (RNNs) | Accuracy (82%) |
| Garraghan et al. [ | Google trace | Study server characteristics | Statistical Analysis | X |
| Pan et al. [ | Google’s | Identify and diagnose | Ganesha’s diagnosis | X |
| Chen Xin et al. [ | Google trace | Study characteristics of failed | Statistical Analysis | X |
| Di Sheng et al. [ | Google trace | Understanding characteristics | K-means | X |
| Lu et al. [ | Alibaba | Understand machine | Statistical Analysis | X |
| Liang et al. [ | Log files from | Predict failure based on | Bursty nature of failure | X |
| Zhang et al. [ | Google’s Compute | Study the problem of deriving | Statistical Analysis | X |
| Di Martino et al. [ | Cloud data | Analyzing causes of SLA | Statistical Analysis | X |
| Chen Weiwei et al. [ | Scientific workflow | Transient failures and | Statistical and | X |
| Samak et al. [ | Scientific workflow | Predicting failure probability | Machine Learning | For Epigenome app |
| Bala and Chana [ | Scientific workflow | Task failure prediction | Machine Learning | NB is the highest accuracy (93%) |
| Rosa et al. [ | Google trace | Failure Prediction | Machine Learning | For ANN, accuracy (76.8%) |
| Liu et al. [ | Google trace | Predicting job | Machine Learning | Accuracy (93.07%) |
| Reiss and et al. [ | Google trace | Highlighting heterogeneous | Statistical Analysis | X |
| Jassas and | Google trace | Study workload features | Statistical Analysis | X |
| Wang et al. [ | Network logs | Focus on risk-aware models | Machine Learning | Accuracy (95%) |
| El-Sayed et al. [ | Google trace | Design a job | Machine Learning | Precision (95%) & Recall (94%) |
| Shetty et al. [ | Google trace | Predicting job termination | Machine Learning | Precision (92%) & Recall (94.8%) |
| Proposed Model | Google trace | Design a failure | Machine Learning | For Google trace |
Figure 1The proposed evaluation process.
Basic description and characteristics of the clusters having traces in the Atlas repository and Google cluster traces.
| Dataset | Nodes | Sample Size | Features | Failed Ratio (%) | Length |
|---|---|---|---|---|---|
| Google Cluster | 12,550 | 28,546,501 | 11 | 36.2 | 29 days |
| LANL Mustang | 1600 | 2,113,175 | 9 | 7.2 | 5 years |
| LANL Trinity | 9408 | 20,277 | 14 | 16.5 | 3 months |
Google trace overview.
| Trace Characteristic | Value |
|---|---|
| Total number of users | 933 users |
| Submitted jobs | 676,975 jobs |
| Scheduling jobs | 676,967 jobs |
| Finished jobs | 386,218 jobs |
| Failed jobs | 10,133 jobs |
| Killed jobs | 272,609 jobs |
| Evict jobs | 22 jobs |
| Lost jobs | 16 jobs |
| Submitted tasks | 48,330,301 tasks |
| Scheduling tasks | 47,306,307 tasks |
| Finished tasks | 18,187,970 tasks |
| Failed tasks | 13,828,583 tasks |
| Killed tasks | 10,337,327 tasks |
| Evict tasks | 5,864,223 tasks |
| Lost tasks | 8754 tasks |
Figure 2A comparison between job and task event failure behaviour in 29 days of Google trace. (a) Failed and finished jobs for 29 days of the Google trace; (b) Failed and finished tasks for 29 days of the Google trace.
Figure 3Number of failed and finished tasks for Google cluster trace, focusing on days 2 and 10. (a) Day 2 of the Google traces; (b) Day 10 of the Google traces.
Figure 4Priority level for failed and finished tasks.
Figure 5Scheduling class for failed and finished tasks.
Figure 6Memory was requested for both unsuccessful and finished tasks.
Figure 7CPU was requested for both unsuccessful and finished tasks.
Figure 8Disk space was requested for both unsuccessful and finished tasks.
Figure 9The average number of tasks requested from 2011 to 2016 for cancelled, failed and finished jobs.
Figure 10The average number of nodes from 2011 to 2016 for cancelled, failed and finished jobs.
Figure 11The average number of nodes for cancelled, failed and finished jobs in the month intervals between 2011 and 2016.
Figure 12Correlation between the execution time and the failed, cancelled and finished jobs.
Figure 13Correlation between Trinity required class types and job status.
Figure 14Correlation between Trinity computing resource types and job status.
Figure 15Distribution of task status for Google traces. (a) Google trace in 7 days; (b) Google trace in 29 days.
Figure 16Distribution of job status for Mustang and Trinity traces. (a) Mustang trace; (b) Trinity trace.
Figure 17Performance evaluation of different algorithms applied to the Google trace. (a) Google trace in 7 days; (b) Google trace in 29 days.
Figure 18Performance evaluation of different machine learning algorithms applied to the Mustang and Trinity traces. (a) Mustang trace; (b) Trinity trace.
Training and testing time and accuracy for all applied ML classifiers.
| DTs | RF | KNN | NB | Gradient Boosting | XGBoost | ||
|---|---|---|---|---|---|---|---|
|
| Training time | 53.8 | 247.6 | 45.7 | 5.2 | 2093.8 | 2180.5 |
| Testing time | 1.03 | 11 | – | 1 | 10.9 | 20.5 | |
| Accuracy | 98 | 98 | – | 78 | 96 | 96 | |
|
| Training time | 13.23 | 75.7 | 5.6 | 0.8 | 344.46 | 182.1 |
| Testing time | 0.16 | 1.9 | 2114.3 | 0.2 | 1.3 | 3.1 | |
| Accuracy | 98 | 98 | 97 | 92 | 97 | 97 | |
|
| Training time | 4.8 | 11.5 | 2.9 | 0.16 | 67.7 | 37.8 |
| Testing time | 0.03 | 0.3 | 5.8 | 0.06 | 0.3 | 0.3 | |
| Accuracy | 99 | 99 | 95 | 86 | 97 | 97 | |
|
| Training time | 0.08 | 0.21 | 0.05 | 0.01 | 0.84 | 0.21 |
| Testing time | 0.0009 | 0.007 | 0.16 | 0.001 | 0.007 | 0.007 | |
| Accuracy | 96 | 96 | 92 | 65 | 94 | 93 |
Figure 19ROC evaluation of different ML algorithms on the Google cluster trace. (a) ROC for the first week of Google trace; (b) ROC for all Google trace observations.
Figure 20ROC evaluation of different ML algorithms on the Mustang and Trinity traces. (a) ROC for Mustang trace; (b) ROC for Trinity trace.
Evaluated results of various feature selection methods.
| Random Forest | Decision Trees | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Prec. | Rec. | F1-score | Train. (t) | Test. (t) | Prec. | Rec. | F1-score | Train. (t) | Test. (t) | ||
|
| SelectKBest | 98% | 97% | 97% | 2683 | 39.5 | 97% | 97% | 97% | 491.3 | 4.17 |
| Feature Importance | 95% | 93% | 94% | 2924 | 60.40 | 93% | 93% | 93% | 518.3 | 7.9 | |
| RFE | 99% | 99% | 99% | 2812 | 36.43 | 99% | 99% | 99% | 542.18 | 3.39 | |
|
| SelectKBest | 72% | 65% | 69% | 0.27 | 0.008 | 72% | 69% | 70% | 0.05 | 0.0008 |
| Feature Importance | 90% | 88% | 90% | 0.38 | 0.02 | 89% | 87% | 88% | 0.07 | 0.001 | |
| RFE | 85% | 81% | 83% | 0.58 | 0.09 | 84% | 83% | 84% | 0.09 | 0.008 | |
|
| SelectKBest | 94% | 94% | 94% | 9.66 | 0.32 | 92% | 93% | 93% | 0.9 | 0.03 |
| Feature Importance | 95% | 94% | 95% | 10.3 | 0.56 | 94% | 94% | 93% | 1.2 | 0.09 | |
| RFE | 94% | 93% | 93% | 11.1 | 0.73 | 93% | 93% | 93% | 1.4 | 0.12 | |