融合预训练表征与伪强标签的声音事件检测方法
DOI:
作者:
作者单位:

北方工业大学 人工智能与计算机学院 北京市 100144

作者简介:

通讯作者:

基金项目:


A method for sound event detection that integrates pre-trained representations with pseudo-strong labels
Author:
Affiliation:

North China University of Technology, Artificial Intelligence and Computer Science, Beijing 100144, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
    摘要:

    声音事件检测(Sound Event Detection, SED)性能提升的关键瓶颈在于对精确时间标注强标签数据的高度依赖。然而,真实场景中采集的音频数据通常呈现强标签、弱标签、无标签及软标签等多源异构监督形式,标注粒度与监督强度差异显著,导致模型难以统一建模并充分挖掘细粒度时序信息。针对上述问题,本文从统一优化角度出发,提出一种融合预训练声学表征与伪强标签学习的半监督声音事件检测方法,将多源异构监督信息建模为显式监督、一致性约束与伪标签监督的联合优化过程。 在模型结构上,引入Transformer预训练声学表征与CNN-RNN检测网络进行多尺度特征融合,以提升模型的全局语义建模能力与局部时频分辨能力;在训练策略上,设计渐进式两阶段微调机制,通过“特征适配—联合优化”的优化路径,实现预训练表征与帧级检测任务之间的稳定对齐;在数据利用层面,提出多模型一致性驱动的伪强标签推断机制,通过预测融合与时间结构重建,将弱标签与无标签数据转化为高质量帧级监督信号,并进一步用于轻量化模型的训练,实现基于数据驱动的知识迁移。 在DCASE 2024 Task 4数据集(DESED、MAESTRO Real及AudioSet Strong)上的实验结果表明,所提出方法能够有效提升模型性能,其中CRNN-ATST模型的PSDS1和mpAUC分别达到0.531和0.740,较基线提升14.8%和9.3%。在引入伪强标签后,轻量化LightSED模型在参数量压缩88.6%的情况下,总体性能仅下降3.46%,且PSDS2等指标与大模型基本持平。结果表明,本文方法在提升异构监督数据利用效率的同时,兼顾检测性能与模型复杂度,为强标签稀缺及资源受限条件下的声音事件检测提供了一种有效解决方案。

    Abstract:

    The key bottleneck hindering performance improvements in Sound Event Detection (SED) lies in its heavy reliance on strongly labelled data with precise temporal annotations. However, audio data collected in real-world scenarios typically presents a multi-source, heterogeneous supervision format comprising strongly labelled, weakly labelled, unlabelled, and soft-labelled data. Significant variations in annotation granularity and supervision strength make it difficult for models to establish a unified representation and fully exploit fine-grained temporal information. To address these issues, this paper proposes a semi-supervised sound event detection method that integrates pre-trained acoustic representations with pseudo-strong labelling learning, from a unified optimisation perspective. It models multi-source heterogeneous supervision as a joint optimisation process involving explicit supervision, consistency constraints, and pseudo-labelling supervision. In terms of model architecture, Transformer-based pre-trained acoustic representations are integrated with a CNN-RNN detection network to achieve multi-scale feature fusion, thereby enhancing the model’s global semantic modelling capabilities and local spatio-temporal resolution. Regarding training strategies, a progressive two-stage fine-tuning mechanism is designed; through an ‘feature adaptation–joint optimisation’ pathway, stable alignment is achieved between the pre-trained representations and the frame-level detection task. In terms of data utilisation, we propose a pseudo-strong label inference mechanism driven by multi-model consistency. Through prediction fusion and temporal structure reconstruction, this mechanism converts weakly labelled and unlabelled data into high-quality frame-level supervision signals, which are subsequently utilised for training lightweight models, thereby achieving data-driven knowledge transfer. Experimental results on the DCASE 2024 Task 4 datasets (DESED, MAESTRO Real and AudioSet Strong) demonstrate that the proposed method effectively enhances model performance. Specifically, the CRNN-ATST model achieves PSDS1 and mpAUC scores of 0.531 and 0.740 respectively, representing improvements of 14.8% and 9.3% over the baseline. Following the introduction of pseudo-strong labels, the lightweight LightSED model exhibits a mere 3.46% decline in overall performance despite an 88.6% reduction in parameters, whilst metrics such as PSDS2 remain on par with those of large models. The results demonstrate that our method balances detection performance with model complexity whilst enhancing the utilisation efficiency of heterogeneous supervised data, thereby providing an effective solution for sound event detection under conditions of scarce strong labels and resource constraints.

    参考文献
    相似文献
    引证文献
引用本文

关晓菡 ,罗伟杰,李嘉锋,蔡希昌.融合预训练表征与伪强标签的声音事件检测方法[J].数据采集与处理,,():

复制
分享
文章指标
  • 点击次数:
  • 下载次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-14