Abstract:Sound Event Detection (SED) plays a critical role in intelligent perception systems, with broad applications in industrial monitoring, ecological surveillance, and public safety. To address the limitations of low-resource SED scenarios, such as insufficient labeled data and limited temporal resolution, this paper proposes a frame-level few-shot sound event detection method based on multi-task learning and dynamic convolution. The proposed method integrates a foreground/background sound classification (FBSC) task alongside the primary SED objective, and combines a NetMamba encoder with a Frequency Dynamic Convolution (FDC) module to enhance long-range temporal modeling and frequency-adaptive feature extraction. Additionally, a linear time-domain perturbation strategy, TimeFilterAug, is designed to simulate noise interference in complex acoustic environments, further improving the model's generalization capability. The proposed method achieved an F1-score of 56.7% in the DCASE 2024 Challenge, ranking second. Experimental results confirm its effectiveness in detecting transient events and maintaining robustness under noisy conditions, demonstrating strong potential for practical deployment.