基于时空多级交互编码的自适应开放词汇脑电解码
作者:
作者单位:

上海理工大学光电信息与计算机工程学院,上海 200093

作者简介:

通讯作者:

基金项目:

上海市自然科学基金(22ZR1443700)。


A Spatio-Temporal Multi-level Adaptive Interactive Attention Network for Open-Vocabulary EEG Decoding
Author:
Affiliation:

School of Optical-Electrical and Computer Engineering, University of Shanghai for Science and Technology, Shanghai 200093, China

Fund Project:

Natural Science Foundation of Shanghai (No.22ZR1443700).

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
    摘要:

    在开放词汇场景下进行脑电信号到文本的解码任务中,从时空关系复杂且具有显著个体差异的脑电信号中提取准确的语义信息是一项关键挑战。针对这一挑战,本文提出了一种基于时空特征增强、受试者自适应机制和多级交互注意力机制的解码方法。首先,结合空洞卷积与门控机制,设计时空特征增强模块,以提升时序动态建模能力和通道依赖性表征能力。随后,引入受试者特定令牌,对不同受试者间的脑电分布差异进行显式建模,缓解跨受试条件下个体差异对模型解码性能的影响。最后,提出多级交互注意力机制,联合通道级、受试者级和全局级注意力,解析脑电信号中的多层次语义关联。此外,首次引入基于可学习激活函数的非线性投影网络,高效表征高维脑电信号,增强模型对脑电信号的解码能力。实验结果表明,本文所提出的方法在词汇相似度指标BLEU和ROUGE以及语义相似度指标BERTScore上分别提升2.23、4.87和4.66个百分点,验证了该方法在开放词汇脑电解码任务中的有效性与鲁棒性。

    Abstract:

    Decoding electroencephalography (EEG) signals into natural language text is a fundamental task in brain-computer interface (BCI) research. Existing approaches often suffer from insufficient modeling of spatiotemporal EEG characteristics, limited capability in capturing hierarchical semantic relationships, and poor generalization across subjects. To address these issues, this study proposes a spatio-temporal multi-level adaptive cross-subject interactive attention network (ST-MSIAN). First, the spatiotemporal feature enhancemant module (SFEM) employs parallel temporal and spatial branches to jointly model temporal dynamics and channel dependencies of EEG signals. Multi-scale dilated convolutions with dilation rates of 1, 2, and 4 are adopted to capture both local and long-range temporal dependencies, while spatial convolutions and attention mechanisms enhance inter-channel representations. Subsequently, subject-specific tokens are introduced to explicitly characterize individual neural patterns and reduce the influence of inter-subject variability. Based on these representations, the multi-level interactive attention mechanism (MIAM) progressively extracts hierarchical semantic information through channel-level, subject-level, and global-level attention interactions. The resulting features are further processed by a six-layer Transformer encoder to model contextual dependencies. Finally, a KAN-based projection head with learnable activation functions is employed to establish a more effective nonlinear mapping between high-dimensional EEG features and the language embedding space of a pretrained BART decoder, enabling accurate text generation. Experiments are conducted on the publicly available Zurich cognitive language processing corpus (ZuCo), including both ZuCo v1.0 and ZuCo v2.0 datasets. The proposed ST-MSIAN achieves a BLEU-1 score of 44.10%, a BLEU-3 score of 16.21%, a ROUGE-F score of 37.62%, and a BERTScore-F score of 60.08%. Compared with the strongest competing method, the proposed approach improves BLEU-1, ROUGE-F, and BERTScore-F by 2.23, 4.87, and 4.66 percentage points, respectively. The proposed ST-MSIAN provides an effective and robust solution for cross-subject neural language decoding and offers promising potential for future applications in assistive communication systems and intelligent brain-computer interfaces.

    参考文献
    相似文献
    引证文献
引用本文

常丽涛,王永雄,黄帅,王哲.基于时空多级交互编码的自适应开放词汇脑电解码[J].数据采集与处理,2026,(4):1010-1025

复制
分享
文章指标
  • 点击次数:
  • 下载次数:
历史
  • 收稿日期:2026-01-23
  • 最后修改日期:2026-07-15
  • 录用日期:
  • 在线发布日期: 2026-08-13