| Literature DB >> 36236270 |
Kaiguo Xia1, Zhisong Pan2, Pengqiang Mao2.
Abstract
Video compression sensing can use a few measurements to obtain the original video by reconstruction algorithms. There is a natural correlation between video frames, and how to exploit this feature becomes the key to improving the reconstruction quality. More and more deep learning-based video compression sensing (VCS) methods are proposed. Some methods overlook interframe information, so they fail to achieve satisfactory reconstruction quality. Some use complex network structures to exploit the interframe information, but it increases the parameters and makes the training process more complicated. To overcome the limitations of existing VCS methods, we propose an efficient end-to-end VCS network, which integrates the measurement and reconstruction into one whole framework. In the measurement part, we train a measurement matrix rather than a pre-prepared random matrix, which fits the video reconstruction task better. An unfolded LSTM network is utilized in the reconstruction part, deeply fusing the intra- and interframe spatial-temporal information. The proposed method has higher reconstruction accuracy than existing video compression sensing networks and even performs well at measurement ratios as low as 0.01.Entities:
Keywords: end-to-end deep learning network; measurement matrix training; unfolded LSTM; video compressing sensing
Mesh:
Year: 2022 PMID: 36236270 PMCID: PMC9572108 DOI: 10.3390/s22197172
Source DB: PubMed Journal: Sensors (Basel) ISSN: 1424-8220 Impact factor: 3.847
Figure 1Principle of SVCS measurement process.
Figure 2The architecture of the proposed framework.
Figure 3The measurement process based on CNN structure.
Figure 4The process of the initial reconstruction.
Figure 5The process of the unfolded LSTM.
Comparison of the proposed method with VCSNet.
| Name | Ratio | Frame Size | GOP |
|
| PSNR | SSIM |
|---|---|---|---|---|---|---|---|
| VCSNet | 0.15 | 96 × 96 | 8 | 0.5 | 0.1 | 34.29 | 0.90 |
| Proposed | 96 × 96 | 8 | 0.5 | 0.1 |
|
| |
| VCSNet | 0.07 | 96 × 96 | 8 | 0.5 | 0.01 | 29.58 | 0.82 |
| Proposed | 96 × 96 | 8 | 0.5 | 0.01 |
|
|
Comparison of the proposed method with CSVideoNet.
| Name | Ratio | Frame Size | GOP |
|
| PSNR | SSIM |
|---|---|---|---|---|---|---|---|
| CSVideoNet | 0.04 | 160 × 160 | 10 | 0.2 | 0.022 | 26.87 | 0.81 |
| Proposed | 160 × 160 | 10 | 0.2 | 0.022 |
|
| |
| CSVideoNet | 0.02 | 160 × 160 | 10 | 0.1 | 0.011 | 25.09 | 0.77 |
| Proposed | 160 × 160 | 10 | 0.1 | 0.011 |
|
| |
| CSVideoNet | 0.01 | 160 × 160 | 10 | 0.06 | 0.004 | 24.23 | 0.74 |
| Proposed | 160 × 160 | 10 | 0.06 | 0.004 |
|
|
Comparison of the proposed method with TVCS methods under ratio 1/16 (0.0625).
| Name | Ratio | Frame Size | GOP |
|
| PSNR | SSIM |
|---|---|---|---|---|---|---|---|
| DCF | 1/16 | 160 × 160 | 16 | - | - | 24.67 | 0.71 |
| C2B | 160 × 160 | 16 | - | - | 32.23 | 0.93 | |
| CSVideoNet | 160 × 160 | 10 | 0.2 | 0.022 | 28.08 | 0.84 | |
| VCSNet | 160 × 160 | 10 | 0.2 | 0.022 | 28.57 | 0.86 | |
| Proposed | 160 × 160 | 10 | 0.2 | 0.022 |
|
|
Figure 6The comparison of reconstructed frames between the proposed and DFC, CSVideo, VCSNet and C2B methods.
Comparison of the proposed method with classical LSTM structure.
| Name | Ratio | Frame Size | GOP |
|
| PSNR | SSIM |
|---|---|---|---|---|---|---|---|
| LSTM | 0.04 | 160 × 160 | 10 | 0.2 | 0.047 | 34.11 | 0.91 |
| Proposed | 160 × 160 | 10 | 0.2 | 0.047 |
|
| |
| LSTM | 0.02 | 160 × 160 | 10 | 0.2 | 0.022 | 31.50 | 0.81 |
| Proposed | 160 × 160 | 10 | 0.2 | 0.022 |
|
| |
| LSTM | 0.01 | 160 × 160 | 10 | 0.1 | 0.011 | 30.32 | 0.77 |
| Proposed | 160 × 160 | 10 | 0.1 | 0.011 |
|
|
Comparison of the proposed method performance in different under the same ratio.
| Frame Size | Patch Size | GOP |
|
|
|
| PSNR | SSIM |
|---|---|---|---|---|---|---|---|---|
| 96 × 96 | 32 × 32 | 10 | 0.5 | 0.014 | 512 | 14 | 31.88 | 0.81 |
| 96 × 96 | 32 × 32 | 10 | 0.4 | 0.025 | 409 | 25 | 33.59 | 0.83 |
| 96 × 96 | 32 × 32 | 10 | 0.3 | 0.036 | 307 | 37 | 34.58 | 0.86 |
| 96 × 96 | 32 × 32 | 10 | 0.2 | 0.047 | 204 | 48 |
|
|
Figure 7Variation curve of model reconstruction performance with the number of training epochs.
Comparison of the time complexity of the existing models and the proposed model.
| Model | CSVideoNet | VCSNet | Proposed |
|---|---|---|---|
| Time(s) | 587.70 | 359.61 |
|