| Literature DB >> 35897985 |
Zhenxiao Zhao1, Lei Zhang2, Huiliang Shang1,2.
Abstract
Falls pose a great danger to social development, especially to the elderly population. When a fall occurs, the body's center of gravity moves from a high position to a low position, and the magnitude of change varies among body parts. Most existing fall recognition methods based on deep learning have not yet considered the differences between the movement and the change in amplitude of each body part. Besides, some problems exist such as complicated design, slow detection speed, and lack of timeliness. To alleviate these problems, a lightweight subgraph-based deep learning method utilizing skeleton information for fall recognition is proposed in this paper. The skeleton information of the human body is extracted by OpenPose, and an end-to-end lightweight subgraph-based network is designed. Sub-graph division and sub-graph attention modules are introduced to add a larger perceptual field while maintaining its lightweight characteristics. A multi-scale temporal convolution module is also designed to extract and fuse multi-scale temporal features, which enriches the feature representation. The proposed method is evaluated on a partial fall dataset collected in NTU and on two public datasets, and outperforms existing methods. It indicates that the proposed method is accurate and lightweight, which means it is suitable for real-time detection and rapid response to falls.Entities:
Keywords: deep learning; fall recognition; skeleton extraction; sub-graph
Mesh:
Year: 2022 PMID: 35897985 PMCID: PMC9332296 DOI: 10.3390/s22155482
Source DB: PubMed Journal: Sensors (Basel) ISSN: 1424-8220 Impact factor: 3.847
Figure 1Schematic diagram of a fall and sub-graph division.
Figure 2Illustration of the proposed method. It consists of skeleton extraction, a spatial module, and a temporal module.
Figure 3Composition of the AGCN. It consists of SGCN, a standard temporal convolution layer, and a sub-graph attention.
Figure 4Illustration of the sub-graph attention module.
Comparison of methods on the UR Fall dataset (%).
| Method | Sensitivity | Specificity | Accuracy |
|---|---|---|---|
| AR-FD [ | 98.0 | 89.4 | 94.0 |
| MEWMA-FD [ | 100 | 94.9 | 96.6 |
| Shi-Tomasi-FD [ | 96.7 | - | 95.7 |
| CNN-FD [ | 100 | 92.0 | 95.0 |
| CNN-LSTM-FD [ | 91.4 | - | - |
| Proposed method | 98.5 | 96.0 | 97.0 |
Comparison of methods on the UP-Fall detection dataset (%).
| Method | Sensitivity | Specificity | Accuracy |
|---|---|---|---|
| CNN + cam1 [ | 97.72 | 81.58 | 95.24 |
| CNN + cam2 [ | 95.57 | 79.67 | 94.78 |
| RF [ | 14.48 | 92.9 | 32.33 |
| SVM [ | 14.30 | 92.97 | 34.40 |
| MLP [ | 10.59 | 92.21 | 27.08 |
| KNN [ | 15.54 | 93.09 | 34.03 |
| CNN [ | 71.3 | 99.5 | 95.1 |
| CNN [ | 99.5 | 83.08 | 95.64 |
| RF + SVM + MLP + KNN [ | 96.80 | 99.11 | 98.59 |
| CNN + LSTM [ | 94.37 | 98.96 | 98.59 |
| Proposed method | 95.43 | 99.12 | 98.85 |
The experimental calculation results on the NTU dataset (six types of actions) (%).
| Sensitivity | Specificity | Accuracy | |
|---|---|---|---|
| Result | 97.5 | 89.6 | 94.5 |
Figure 5Confusion matrix for falls and similar actions on the collected NTU dataset.
Performance of the proposed method on the collected NTU dataset (%).
| Setting | NTU | |
|---|---|---|
| X-Sub | X-View | |
| Sub-Graph Division | 90.3 | 92.8 |
| Sub-Graph Attention | 90.2 | 92.5 |
| Sub-Graph Division + Sub-Graph Attention |
|
|
Comparison of temporal convolution with different scales (%).
| Kernel Size | X-Sub | ||
|---|---|---|---|
| 3 | 5 | 7 | |
| 85.7 | |||
| 87.5 | |||
| √ | √ | 87.6 | |
| √ | 87.2 | ||
| √ | √ | 88.9 | |
| √ | √ | 88.4 | |
| √ | √ | 88.7 | |
| √ | √ | √ | 89.8 |
Figure 6Sub-graph attention visualization feature map of a fall event.
Time and memory cost analysis on the three datasets.
| Dataset | Training Time (h) | Speed (fp/s) |
|---|---|---|
| collected NTU dataset | 2 | 31 |
| URFD | 4.5 | 32 |
| UP-Fall | 7.5 | 30 |