| Literature DB >> 34960499 |
Fuad Al Abir1, Md Al Siam1, Abu Sayeed1, Md Al Mehedi Hasan1,2, Jungpil Shin2.
Abstract
The act of writing letters or words in free space with body movements is known as air-writing. Air-writing recognition is a special case of gesture recognition in which gestures correspond to characters and digits written in the air. Air-writing, unlike general gestures, does not require the memorization of predefined special gesture patterns. Rather, it is sensitive to the subject and language of interest. Traditional air-writing requires an extra device containing sensor(s), while the wide adoption of smart-bands eliminates the requirement of the extra device. Therefore, air-writing recognition systems are becoming more flexible day by day. However, the variability of signal duration is a key problem in developing an air-writing recognition model. Inconsistent signal duration is obvious due to the nature of the writing and data-recording process. To make the signals consistent in length, researchers attempted various strategies including padding and truncating, but these procedures result in significant data loss. Interpolation is a statistical technique that can be employed for time-series signals to ensure minimum data loss. In this paper, we extensively investigated different interpolation techniques on seven publicly available air-writing datasets and developed a method to recognize air-written characters using a 2D-CNN model. In both user-dependent and user-independent principles, our method outperformed all the state-of-the-art methods by a clear margin for all datasets.Entities:
Keywords: air-writing recognition; convolutional neural network; human–computer interaction; interpolation; time-series data
Mesh:
Year: 2021 PMID: 34960499 PMCID: PMC8705512 DOI: 10.3390/s21248407
Source DB: PubMed Journal: Sensors (Basel) ISSN: 1424-8220 Impact factor: 3.576
Summary of the datasets used in this study.
| Dataset | No. of Classes, | No. of Features, | No. of Users | No. of Samples | Training Principle | |
|---|---|---|---|---|---|---|
| User Dependent | User Independent | |||||
| RTD * [ | 10 | 2 | 10 | 20,000 | ✓ | ✗ |
| RTC * [ | 26 | 3 | 10 | 30,000 | ✓ | ✗ |
| Smart-band [ | 26 | 6 | 55 | 21,450 | ✓ | ✓ |
| 6DMG-digit [ | 10 | 13 | 6 | 600 | ✓ | ✓ |
| 6DMG-lower [ | 26 | 13 | 6 | 1470 | ✓ | ✓ |
| 6DMG-upper [ | 26 | 13 | 25 | 6500 | ✓ | ✓ |
| 6DMG-all [ | 62 | 13 | 25 | 8570 | ✓ | ✓ |
* The authors did not disclose user data. Therefore, training in user-independent principle cannot be performed.
Figure 1Different interpolation techniques applied on time-series data for upsampling and downsampling. (A) for 1/4 downsampling (signal length from 100 to 25), (B) for 1/2 downsampling (signal length from 100 to 50), (C) for 2× upsampling (signal length from 100 to 200), (D) for 4× upsampling (signal length from 100 to 400). The input signal was taken from the smart-band dataset.
Network architecture of the 2D-CNN model for air-writing recognition based on smart-band dataset.
| Operation Group | Layer Name | Filter Size | No. of Filters | Stride Size | Padding Size | Activation Function | Output Size * | No. of Parameters * |
|---|---|---|---|---|---|---|---|---|
| - | Input | - | - | - | - | - |
| 0 |
| Group1 | Conv1-1 |
| 32 |
|
| ReLU |
| 160 |
| Conv1-2 |
| 32 |
|
| ReLU |
| 4128 | |
| MaxPool1 |
| 1 |
| 0 | - |
| 0 | |
| Dropout |
|
| 0 | |||||
| Group2 | Conv2-1 |
| 64 |
|
| ReLU |
| 8256 |
| Conv2-2 |
| 64 |
|
| ReLU |
| 16,448 | |
| MaxPool2 |
| 1 |
| 0 | - |
| 0 | |
| Dropout |
|
| 0 | |||||
| Group3 | Conv3-1 |
| 128 |
|
| ReLU |
| 32,896 |
| Conv3-2 |
| 128 |
|
| ReLU |
| 65,664 | |
| MaxPool3 |
| 1 |
| 0 | - |
| 0 | |
| Dropout |
|
| 0 | |||||
| Group4 | Flatten | - | - | - | - | - | 3200 | 0 |
| Dense | - | - | - | - | ReLU | 512 | 1,638,912 | |
| Dropout |
| 512 | 0 | |||||
| Dense | - | - | - | - | Softmax | 26 | 13,338 | |
| Total | 1,779,802 | |||||||
* Output size and no. of parameters vary based on the number of features and signal length, l, depending upon the dataset under consideration. For smart-band dataset, the number of features is 6 and the signal length, l is 200 (see Table 1 and Section 3.2.1). Therefore, we yield this 2D-CNN network. The layers that construct the network and the attributes remain the same for all datasets.
Figure 2Histogram of the length of all samples from smart-band dataset.
Interpolation methods in different upsampling and downsampling settings.
| Upsampling | Downsampling | Signal Length, | Accuracy | |
|---|---|---|---|---|
| Avg. (%) | Std. | |||
| Bicubic | Bicubic | 100 | 87.35 | 0.41 |
| Lanczos | 87.21 | 0.59 | ||
| Bilinear | 87.76 | 0.45 | ||
| Nearest neighbor | 87.04 | 0.21 | ||
| Lanczos | Bicubic | 100 | 87.54 | 0.18 |
| Lanczos | 86.50 | 0.70 | ||
| Bilinear | 86.84 | 0.46 | ||
| Nearest neighbor | 86.73 | 0.27 | ||
| Bilinear | Bicubic | 100 | 86.91 | 0.77 |
| Lanczos | 86.77 | 0.63 | ||
| Bilinear | 86.25 | 0.80 | ||
| Nearest neighbor | 87.38 | 0.09 | ||
| Nearest neighbor | Bicubic | 100 | 87.38 | 0.37 |
| Lanczos | 87.32 | 0.76 | ||
| Bilinear | 86.67 | 0.58 | ||
| Nearest neighbor | 87.08 | 1.24 | ||
| Bicubic | Bicubic | 200 | 88.54 | 0.31 |
| Lanczos | Lanczos | 87.35 | 0.31 | |
| Bilinear | Bilinear | 88.46 | 0.19 | |
| Nearest neighbor | Nearest neighbor | 88.08 | 0.59 | |
Comparative analysis of Bicubic interpolation with padding and truncation methods.
| Approach | Sequence | # Padded or | # Truncated or | # Flops | Inference | Accuracy | |
|---|---|---|---|---|---|---|---|
| Avg. (%) | Std. | ||||||
| Pre-sequence | 50 | 212 | 21,238 | 599,177 | 1.648 | 59.87 | 0.52 |
| 100 | 10,892 | 10,558 | 992,393 | 1.492 | 84.57 | 0.30 | |
| 200 | 21,161 | 289 | 1,778,825 | 1.709 | 86.58 | 0.37 | |
| 400 | 21,449 | 1 | 3,417,225 | 2.056 | 86.38 | 0.24 | |
| Post-sequence | 50 | 212 | 21,238 | 599,177 | 1.429 | 48.15 | 0.44 |
| 100 | 10,892 | 10,558 | 992,393 | 1.843 | 80.25 | 0.23 | |
| 200 | 21,161 | 289 | 1,778,825 | 1.770 | 85.84 | 0.42 | |
| 400 | 21,449 | 1 | 3,417,225 | 2.519 | 85.63 | 0.35 | |
| Bicubic | 50 | 212 | 21,238 | 599,177 | 1.475 | 84.98 | 0.35 |
| 100 | 10,892 | 10,558 | 992,393 | 1.533 | 87.35 | 0.41 | |
| 200 | 21,161 | 289 | 1,778,825 | 1.687 | 88.54 | 0.31 | |
| 400 | 21,449 | 1 | 3,417,225 | 2.430 | 88.02 | 0.48 | |
Selected signal length, l for all datasets.
| Dataset | Min | Max | Signal Length, |
|---|---|---|---|
| RTD | 18 | 150 | 125 |
| RTC | 21 | 173 | 125 |
| Smart-band | 34 | 438 | 200 |
| 6DMG-digit | 29 | 218 | 175 |
| 6DMG-lower | 27 | 163 | 150 |
| 6DMG-upper | 27 | 412 | 250 |
| 6DMG-all | 27 | 412 | 250 |
The terms “Min” and “Max” represent the minimum and maximum length of the signals in that particular dataset, respectively.
Performance evaluation for user-dependent and independent method on smart-band and RealSense-based Trajectory datasets. Abbreviations of the approaches are given in the Abbreviations section of this paper.
| Training Principle | Approach | Accuracy | ||
|---|---|---|---|---|
| Smart-Band | RTC | RTD | ||
| User-dependent | KNN-DTW based [ | 89.20 | - | - |
| 2D-CNN based [ | - | 97.29 | - | |
| LSTM based [ | - | - | 99.17 | |
| CNN-LSTM fusion [ | - | 98.74 | 99.63 | |
| Proposed | 91.34 | 99.63 | 99.76 | |
| User-independent | 1D-CNN [ | 83.20 | - | - |
| Proposed | 85.59 | - | - | |
Performance evaluation for user-dependent and independent methods on the 6DMG dataset. Abbreviations of the approaches are given in the Abbreviations section of this paper.
| Training | Approach | Accuracy | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Digit | Lower | Upper | All | ||||||
| Avg. (%) | Std. | Avg. (%) | Std. | Avg. (%) | Std. | Avg. (%) | Std. | ||
| User- | HMM-based [ | - | - | - | - | 98.16 | 2.37 | - | - |
| LSTM-bases [ | 97.33 | 1.49 | 96.80 | 0.57 | 98.34 | 0.50 | 94.75 | 0.31 | |
| CRF-CNN fusion [ | - | - | - | - | 98.57 | - | - | - | |
| BiLSTM-CNN fusion [ | 99.33 | - | - | - | 99.27 | - | - | - | |
| CHMM-based [ | 99.00 | 1.09 | 98.22 | 0.73 | 97.29 | 0.66 | 95.91 | 0.47 | |
| UDA [ | 99.78 | 0.03 | 98.94 | 0.08 | 99.55 | 0.06 | 97.03 | 0.11 | |
| Proposed | 100.00 | 0.00 | 99.47 | 0.39 | 99.80 | 0.20 | 98.99 | 0.23 | |
| User- | CHMM [ | 96.70 | 4.08 | 76.38 | 5.25 | 91.03 | 1.54 | 62.69 | 2.91 |
| UDA [ | 98.74 | 0.34 | 92.86 | 0.48 | 96.99 | 0.45 | 87.69 | 0.58 | |
| Proposed | 99.26 | 0.12 | 94.48 | 0.45 | 99.23 | 0.94 | 91.24 | 0.86 | |