A Multi-Scale Temporal Cross-Attention Network for Video Compressive Sensing Reconstruction
DOI:
CSTR:
Author:
Affiliation:

School of Internet of Things, Nanjing University of Posts and Telecommunications, Nanjing 210003, China

Clc Number:

Fund Project:

  • Article
  • |
  • Figures
  • |
  • Metrics
  • |
  • Reference
  • |
  • Related
  • |
  • Cited by
  • |
  • Materials
  • |
  • Comments
    Abstract:

    Video compressive sensing (VCS) provides an efficient acquisition paradigm by sampling video signals at rates far below the Nyquist requirement, but severe undersampling also makes video reconstruction a highly ill-posed inverse problem. To tackle this ill-posed reconstruction problem, this paper proposes a multi-scale temporal cross-attention network for VCS reconstruction, which improves reconstruction quality by exploiting the spatiotemporal redundancy within video sequences. The proposed method adopts an end-to-end reconstruction framework composed of a learnable compressive sampling module and a temporal feature propagation network. At the sampling stage, a learnable block-based sensing module is used to replace the fixed random measurement matrix, so that the sampling operator can be jointly optimized with the reconstruction network. At the reconstruction stage, the measurements are first mapped to an initial estimate of the video frames. Based on this initial reconstruction, bidirectional temporal propagation is further employed to introduce information from neighboring frames and refine the current frame representation. Within each propagation direction, a temporal cross-attention mechanism is designed to model the correlation between the current frame and its adjacent frame. Different from conventional self-attention, the query is generated from the current frame, while the key and value are constructed from both the current and neighboring frames. This design enables the network to retrieve temporal information from adjacent frames and integrate it into the reconstruction of the current frame. Considering that motion patterns vary significantly across different video scenes, multi-scale content representation is introduced into the value branch of the cross-attention module. This design enhances the ability of the network to represent local details and broader motion-related structures without disturbing the correlation calculation in the query and key branches. To reduce the computational burden caused by long video sequences and high-resolution features, linear attention is adopted instead of standard attention. Furthermore, a gated convolutional feed-forward module is introduced after temporal feature fusion. By combining local convolutional modeling with adaptive gating, this module selectively enhances reliable propagated features and suppresses unreliable updates, which helps alleviate error accumulation during recursive reconstruction.Experimental results under multiple sampling rates demonstrate the effectiveness of the proposed method. Compared with existing reconstruction approaches, the proposed network achieves improvements of 2.55 dB in PSNR and 0.10 in SSIM at the sampling rate of 0.01, and 0.93 dB in PSNR and 0.02 in SSIM at the sampling rate of 0.10. These results indicate that the proposed framework can make better use of inter-frame information, especially under severe undersampling conditions. Overall, the proposed method provides an effective reconstruction solution for video compressive sensing and shows strong potential for low-power and resource-limited video acquisition applications.

    Reference
    Related
    Cited by
Get Citation

ZHOU Chao, CHEN Can, ZHANG Dengyin. A Multi-Scale Temporal Cross-Attention Network for Video Compressive Sensing Reconstruction[J]. Journal of Data Acquisition and Processing,,().

Copy
Related Videos

Share
Article Metrics
  • Abstract:
  • PDF:
  • HTML:
  • Cited by:
History
  • Received:
  • Revised:
  • Adopted:
  • Online: July 14,2026
  • Published:
Article QR Code