基于多尺度时域交叉注意力的视频压缩感知重构网络
DOI:
作者:
作者单位:

南京邮电大学物联网学院,南京 210003

作者简介:

通讯作者:

基金项目:


A Multi-Scale Temporal Cross-Attention Network for Video Compressive Sensing Reconstruction
Author:
Affiliation:

School of Internet of Things, Nanjing University of Posts and Telecommunications, Nanjing 210003, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
    摘要:

    视频压缩感知能够在编码端以低于奈奎斯特采样定律要求的采样率获取测量值,从而显著降低采集与编码的计算复杂度和功耗,在资源受限的视频采集与传输场景中具有重要应用价值。然而,欠采样导致的测量信息不足使重构问题呈现典型的病态特性,难以仅依赖单帧信息恢复出高质量视频。为了解决此问题,本文提出一种基于多尺度时域交叉注意力的视频压缩感知重构网络,通过充分挖掘视频信号的时空冗余实现高质量重构。针对不同场景下运动幅度与作用范围存在的差异性,本文在注意力融合阶段引入并行多尺度卷积表征,以兼顾不同尺度下的时空信息整合;为控制长序列带来的计算开销,网络采用线性注意力降低计算复杂度并提升推理效率;为了抑制递归重构过程中的不可靠信息并缓解误差随时间累积扩散,本文采用一种门控卷积前馈模块,将局部卷积建模与自适应门控机制结合,对跨帧传播的特征进行选择性增强与抑制。大量实验结果表明,与现有方法相比,本文所提出方法在多种采样率下都取得了更优异的重构性能。

    Abstract:

    Video compressive sensing (VCS) provides an efficient acquisition paradigm by sampling video signals at rates far below the Nyquist requirement, but severe undersampling also makes video reconstruction a highly ill-posed inverse problem. To tackle this ill-posed reconstruction problem, this paper proposes a multi-scale temporal cross-attention network for VCS reconstruction, which improves reconstruction quality by exploiting the spatiotemporal redundancy within video sequences. The proposed method adopts an end-to-end reconstruction framework composed of a learnable compressive sampling module and a temporal feature propagation network. At the sampling stage, a learnable block-based sensing module is used to replace the fixed random measurement matrix, so that the sampling operator can be jointly optimized with the reconstruction network. At the reconstruction stage, the measurements are first mapped to an initial estimate of the video frames. Based on this initial reconstruction, bidirectional temporal propagation is further employed to introduce information from neighboring frames and refine the current frame representation. Within each propagation direction, a temporal cross-attention mechanism is designed to model the correlation between the current frame and its adjacent frame. Different from conventional self-attention, the query is generated from the current frame, while the key and value are constructed from both the current and neighboring frames. This design enables the network to retrieve temporal information from adjacent frames and integrate it into the reconstruction of the current frame. Considering that motion patterns vary significantly across different video scenes, multi-scale content representation is introduced into the value branch of the cross-attention module. This design enhances the ability of the network to represent local details and broader motion-related structures without disturbing the correlation calculation in the query and key branches. To reduce the computational burden caused by long video sequences and high-resolution features, linear attention is adopted instead of standard attention. Furthermore, a gated convolutional feed-forward module is introduced after temporal feature fusion. By combining local convolutional modeling with adaptive gating, this module selectively enhances reliable propagated features and suppresses unreliable updates, which helps alleviate error accumulation during recursive reconstruction.Experimental results under multiple sampling rates demonstrate the effectiveness of the proposed method. Compared with existing reconstruction approaches, the proposed network achieves improvements of 2.55 dB in PSNR and 0.10 in SSIM at the sampling rate of 0.01, and 0.93 dB in PSNR and 0.02 in SSIM at the sampling rate of 0.10. These results indicate that the proposed framework can make better use of inter-frame information, especially under severe undersampling conditions. Overall, the proposed method provides an effective reconstruction solution for video compressive sensing and shows strong potential for low-power and resource-limited video acquisition applications.

    参考文献
    相似文献
    引证文献
引用本文

周超,陈灿,张登银.基于多尺度时域交叉注意力的视频压缩感知重构网络[J].数据采集与处理,,():

复制
分享
文章指标
  • 点击次数:
  • 下载次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-14