Abstract:The steady-state visual evoked potential (SSVEP) technology is rapidly developing in the field of brain-computer interfaces. This paper targets a four-class frequency recognition task to address the low recognition rate caused by insufficient feature extraction and the influence of noise and other non-target signals in short time windows. To this end, a convolutional neural network for EEG signal classification integrating a sparse attention mechanism with time–frequency dual-domain features (SATF-CNN) is proposed. Firstly, a dynamic sinusoidal position encoding module is embedded at the input end. Then, a channel extraction path with a Top-k hybrid sparse attention mechanism is constructed, and key features are selected through multi-threshold sparse fraction selection. Finally, a dual-branch feature extraction network is built, with a time-domain convolutional network branch and a fast Fourier transform frequency-domain network branch capturing time-domain and frequency-domain features respectively. The Kolmogorov-Arnold network is used to achieve nonlinear feature fusion. Experimental results show that in a four-class cross-subject experiment with a 1-second time window, the accuracy rate is as high as 93.54%, and the information transmission rate is 93.13 bit/min. The SATF-CNN model demonstrates outstanding classification performance in SSVEP recognition tasks, providing a new solution for clinical neuroengineering and the development of intelligent interaction devices.