Abstract:Due to factors such as medical ethics, recording conditions, participant cooperation, and annotation costs, high-quality Parkinson's disease (PD) speech annotation data is relatively scarce, which to some extent restricts the application of supervised learning methods in speech-based PD detection. Although the above problems can be overcome through cross-lingual detection and the use of multi-types of corpora, most existing studies have adopted homogeneous corpus alignment strategies, ignoring the complexity and diversity of speech expression in natural contexts, making it difficult to fully extract the deep pathological features of PD. To this end, a continuous self-supervised cross-lingual PD detection method is proposed in this paper, aiming to fully utilize speech data resources for PD detection while avoiding time-consuming and laborious manual annotation. First, based on multi-type speech data in the source language, a sequential fine-tuning strategy is adopted to adaptively train the top-level of the self-supervised feature extraction model. To prevent corpus conflicts and forgetting issues, a continuous learning replay buffer and feature distillation mechanism are introduced to enhance the plasticity of the self-supervised feature extraction model while improving its ability to retain old knowledge. Finally, the model is transferred to the target language and the ability of self-supervised models to model "language independent" features is utilized to achieve effective recognition of PD in cross-lingual scenarios. The experimental results on Italian and Chinese PD datasets show that the proposed method extracts more discriminative features and has better classification accuracy and generalization ability than existing methods in cross-lingual PD detection.