基于多任务学习与动态卷积的帧级别小样本声音事件检测方法
DOI:
作者:
作者单位:

1.科大讯飞股份有限公司讯飞研究院,合肥 230088;2.中国科学技术大学信息科学技术学院,合肥230026 ; 3.中国矿业大学信息与控制工程学院,徐州 221116

作者简介:

通讯作者:

基金项目:


Frame-Level Few-Shot Sound Event Detection Framework Based on Multi-Task Learning and Dynamic Convolution
Author:
Affiliation:

1.iFlytek Research, iFLYTEK Co., Ltd., Hefei 230088, China; 2.School of Information Science and Technology, University of Science and Technology of China, Hefei 230026, China; 3.School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
    摘要:

    声音事件检测(Sound Event Detection, SED)作为智能感知与环境交互的关键技术,广泛应用于工业监测、生态保护和安防预警等领域。为应对SED任务中标注样本稀缺、事件时间分辨率不足等问题,本文提出了一种融合多任务学习与动态卷积的帧级别小样本声音事件检测方法。该方法在帧级别检测模型的基础上引入前景/背景声音二分类辅助任务,结合基于状态空间建模的NetMamba编码器与频率动态卷积结构(Frequency Dynamic Convolution, FDC)提升模型的长时序建模能力与频域特征自适应提取能力。同时,设计了一种线性时域扰动的数据增强策略TimeFilterAug,以模拟复杂声学环境下的噪声干扰,增强模型的泛化能力。所提方法在DCASE 2024国际挑战赛中取得56.7%的F1分数,排名第二。实验结果表明,该框架在短时瞬态事件识别及噪声鲁棒性方面具有显著优势,具备良好的工程应用前景。

    Abstract:

    Sound Event Detection (SED) plays a critical role in intelligent perception systems, with broad applications in industrial monitoring, ecological surveillance, and public safety. To address the limitations of low-resource SED scenarios, such as insufficient labeled data and limited temporal resolution, this paper proposes a frame-level few-shot sound event detection method based on multi-task learning and dynamic convolution. The proposed method integrates a foreground/background sound classification (FBSC) task alongside the primary SED objective, and combines a NetMamba encoder with a Frequency Dynamic Convolution (FDC) module to enhance long-range temporal modeling and frequency-adaptive feature extraction. Additionally, a linear time-domain perturbation strategy, TimeFilterAug, is designed to simulate noise interference in complex acoustic environments, further improving the model's generalization capability. The proposed method achieved an F1-score of 56.7% in the DCASE 2024 Challenge, ranking second. Experimental results confirm its effectiveness in detecting transient events and maintaining robustness under noisy conditions, demonstrating strong potential for practical deployment.

    参考文献
    相似文献
    引证文献
引用本文

方昕 ,赵鹏源,张雨涛,严根伟 ,闫祖龙,邹亮.基于多任务学习与动态卷积的帧级别小样本声音事件检测方法[J].数据采集与处理,,():

复制
分享
文章指标
  • 点击次数:
  • 下载次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-14