Information

Responsible Institution:China Association for Science and Technology

Sponsored by:Chinese Institute of Electronics
Nanjing University of Aeronautics and Astronautics

ISSN:1004-9037

CN:32-1367/TN

Address:29Yudao Street,Nanjing,China

Telephone:025-84892742

Chief Editor:025-84892742

E-mail:sjcj@nuaa.edu.cn

Post Coder:210016

:China Association for Science and Technology
About periodical
  • ISSN 1004-9037
    CN 32-1367/TN
    Abstracting and Indexing
    · Chinese Core Periodicals
    · Chinese Science Citation Database(CSCD)
    · Chinese Core Journals of Science and Technology
    · China National Knowledge Infrastructure(CNKI)
    · China Science and Technology Journal Database (VIP)
    · Chinese Core Journal Database (Wanfang Data)
    · Chinese Academic Journal Comprehensive Evaluation Database(CAJCED)
    · Scopus
    · EBSCO
    · Directory of Open Access Journals (DOAJ)
    · INSPEC

Current Issue
Online First
  • ZHENG Weihao, DU Yafeng, FU Yu, ZHANG Zhe

    2026(4):910-945, DOI: 10.16337/j.1004-9037.2026.04.002

    Abstract:

    Brain-computer interface (BCI) establishes information pathways between the human brain and external devices by analyzing brain activity, providing new technical means for assisted communication, functional assessment, rehabilitation training, and neuromodulation. Electroencephalography (EEG) is one of the most widely used signal modalities in non-invasive BCI and brain health research because of its non-invasiveness, low cost, high temporal resolution, and ease of repeated acquisition. Driven by recent advances in artificial intelligence, EEG decoding has expanded from conventional control command recognition to complex brain-state representation, visual and linguistic semantic reconstruction, and real-time closed-loop interaction; accordingly, this paper reviews representative progress in control intent decoding, brain-state decoding, high-level semantic decoding, general representation learning, and closed-loop interaction, and analyzes their applications in stroke rehabilitation, language and visual function rehabilitation, epilepsy detection and early warning, assessment of disorders of consciousness, and auxiliary diagnosis of mental disorders. Although the representational capacity and application scope of EEG decoding continue to expand, clinical maturity varies markedly across different directions: control intent recognition, abnormal-state detection, and closed-loop neurorehabilitation have accumulated a certain research foundation, whereas visual and linguistic semantic decoding remains largely limited to healthy participants and controlled experimental settings; major barriers to clinical translation include cross-individual and cross-center generalization, neural information fidelity, long-term stability, incremental clinical value, closed-loop safety, and ethical regulation. Future research should be guided by real clinical needs and patient outcomes to advance intelligent EEG decoding from methodological feasibility toward stable, reliable, and evaluable clinical applications.

  • LIANG Zhen, ZHONG Shuting, XU Ying, HUANG Yi, ZHANG Li, HUANG Gan, ZHOU Yongjie

    2026(4):946-965, DOI: 10.16337/j.1004-9037.2026.04.003

    Abstract:

    Non-suicidal self-injury (NSSI) is a common high-risk behavior among adolescents and is closely associated with emotion regulation difficulties, increased suicide risk, and a range of psychiatric disorders. In recent years, advances in magnetic resonance imaging (MRI) have provided important tools for elucidating the neurobiological mechanisms underlying adolescent NSSI. This review summarizes neuroimaging evidence related to adolescent NSSI, with a focus on structural magnetic resonance imaging (sMRI) and functional magnetic resonance imaging (fMRI) studies. Existing evidence suggests that adolescents with NSSI show gray matter structural abnormalities and impaired white matter integrity in the anterior cingulate cortex, insula, frontal and temporal regions, as well as reward-related brain areas. Resting-state and task-based fMRI studies have further identified abnormal functional connectivity and neural activation in brain networks associated with emotion processing, pain regulation, reward processing, social cognition, and executive control. In addition, multimodal studies suggest that disrupted structure-function coupling may represent an important neural mechanism of NSSI. However, current research is limited by relatively small sample sizes, the predominance of cross-sectional designs, difficulties in controlling for psychiatric comorbidities, and substantial heterogeneity in analytical methods. Future studies should employ large-scale, multicenter, longitudinal, and multimodal approaches to identify stable and reliable neuroimaging biomarkers, thereby providing a neurobiological basis for the early identification, risk assessment, and precision intervention of adolescent NSSI.

  • YING Yiyang, QIU Shengliang, CHEN Xi’ang, JIANG Haiteng

    2026(4):966-980, DOI: 10.16337/j.1004-9037.2026.04.004

    Abstract:

    Major depressive disorder (MDD) is a prevalent psychiatric disorder. Conventional screening and diagnosis mainly rely on structured interviews and rating scales, whose subjectivity and limited accessibility hinder early identification in primary care and community settings. To address this problem, this study develops and validates P-MFCNet, a portable multimodal framework for depression recognition based on single-channel electroencephalography (EEG) and wearable electrocardiography (ECG). A multi-task protocol comprising eyes-open rest, negative emotion elicitation, eyes-closed rest, and a semi-structured clinical interview was designed, and synchronized multimodal data were collected at two medical centers. P-MFCNet employs a hierarchical fusion strategy based on Transformer encoders and classification tokens (CLS) to model complementary information across modalities and differences among task states. Modality-level and task-level representations are learned separately and integrated through hierarchical CLS aggregation to generate a subject-level representation for MDD-versus-HC classification using a ResNet head. In single-center five-fold cross-validation, P-MFCNet achieves an accuracy of 0.842 4 and an AUC of 0.796 7. In cross-center external testing, it achieves an accuracy of 0.810 0 and an AUC of 0.875 0. The model outperforms unimodal EEG and ECG models in accuracy, sensitivity, and overall performance, while maintaining a favorable balance between sensitivity and specificity compared with baseline methods. Post-hoc analyses using zero-out ablation and attention visualization show that ECG/HRV features provide discriminative evidence, whereas EEG contributes complementary information under negative emotion elicitation and resting conditions. These findings demonstrate the feasibility of portable EEG-ECG fusion for depression screening and indicate cross-center generalizability and interpretability.Highlights:1. A portable single-channel EEG and wearable ECG framework is developed for scalable and accessible MDD screening.2. Hierarchical CLS-based fusion effectively integrates complementary modality information and task-specific physiological responses.3. P-MFCNet achieves strong cross-center performance and provides interpretable evidence from ECG/HRV and EEG features.

  • JIANG Lin, SONG Xipeng, MA Shuqi, WANG Guangying, YAO Dezhong, XU Peng, LU Jing, LI Fali

    2026(4):981-994, DOI: 10.16337/j.1004-9037.2026.04.005

    Abstract:

    Motor imagery (MI) is a key cognitive process in brain-computer interfaces and motor function rehabilitation, yet substantial individual differences exist in its behavioral performance. Previous studies, largely based on unimodal data, have struggled to reveal the brain network organization mechanisms with high spatiotemporal resolution that are closely associated with behavioral performance during MI. To address this gap, this study constructs a multimodal covariant network (MCN) based on electroencephalography (EEG) and functional magnetic resonance imaging (fMRI) data. We focus on analyzing the topological differences of MCN across three types of MI tasks, its predictive performance regarding individual imagery ability, and the differential network mechanisms between participants with high versus low task imagery ability, aiming to establish the mapping between EEG-fMRI spatiotemporal fusion features and neural representation efficacy of MI. The results show that the MCNs under all three action conditions exhibit clear task-dependent characteristics, with both left-hand and right-hand tasks demonstrating richer and more stable cross-subnetwork connections compared to the foot task, and the behavior-related network being most extensive for the right-hand task. Furthermore, based on the questionnaire scores, participants were divided into high and low MI ability groups. Between-group statistical differences reveal significant topological differences in MCN between the two groups across all three tasks, mainly distributed in the default mode network, sensorimotor network, visual network, and frontoparietal network. These findings demonstrate that MCN effectively captures the task specificity and individual differences in MI, providing new evidence for establishing the mapping between EEG-fMRI spatiotemporal information and the neural representation efficacy of MI from the perspective of multimodal fusion networks.Highlights:1. Propose a multimodal covariant network (MCN) framework that integrates EEG phase-locking value and fMRI beta-series correlation into a symmetric fusion architecture, enabling the unified characterization of cross-modal brain network organization during motor imagery (MI) without presupposing a dominant modality.2. Construct task-specific MCNs for three MI tasks, and systematically reveal hierarchical and task-dependent connectivity patterns, with hand tasks, particularly the right hand, exhibiting richer and more stable cross-subnetwork interactions than foot tasks.3. Design a nested leave-one-out cross-validation prediction pipeline combining least absolute shrinkage and selection operator (LASSO) regression and connectome-based predictive modeling (CPM), demonstrating that MCN features can predict both subjective imagery ability (KVIQ-KI scores) and objective MI classification accuracy, with right-hand MI achieving the strongest predictive performance.

  • ZHANG Xuejun, DONG Xuan

    2026(4):995-1009, DOI: 10.16337/j.1004-9037.2026.04.006

    Abstract:

    Most motor imagery-based brain-computer interface studies rely on single-scale feature extraction methods, which use convolutional or recurrent structures with fixed receptive fields and struggle to comprehensively capture the temporal characteristics of electroencephalogram (EEG) signals. Aiming at the above problem, we propose a multi-scale dilated convolutional network (MSDCN) model for motor imagery EEG signal classification. The proposed model first extracts spatiotemporal features through two layers of one-dimensional convolution, then enhances its ability to model both short-term dynamic variations and long-term dependencies using multi-scale dilated convolutions. Additionally, a squeeze-and-excitation (SE) module is integrated to learn channel-wise feature importance, highlighting key feature channels and improving classification performance. Experimental results show that on the BCI Competition Ⅳ 2a dataset, the model achieves an accuracy of 84.1% for within-subject experiments and 69.1% for cross-subject experiments, and on the BCI Competition Ⅳ 2b dataset, the within-subject accuracy reaches 89.8%.Highlights:1. A novel multi-scale dilated convolutional network (MSDCN) is proposed for motor imagery EEG classification, which adopts four parallel branches with different convolution kernel sizes (1×3, 1×5, 1×7, 1×9) to synchronously capture short-term local temporal features and long-range global dependencies of EEG signals.2. A descending dilation rate strategy (d=8, 4, 2) is designed inside each dilated convolution block, which models global temporal information first and then supplements local details, effectively alleviating the sparse feature loss problem caused by large dilation rates.3. The squeeze-and-excitation (SE) channel attention module is embedded in each multi-scale branch to adaptively assign weights to feature channels, strengthening discriminative EEG channels related to motor imagery tasks and suppressing irrelevant interference features.

  • CHANG Litao, WANG Yongxiong, HUANG Shuai, WANG Zhe

    2026(4):1010-1025, DOI: 10.16337/j.1004-9037.2026.04.007

    Abstract:

    Decoding electroencephalography (EEG) signals into natural language text is a fundamental task in brain-computer interface (BCI) research. Existing approaches often suffer from insufficient modeling of spatiotemporal EEG characteristics, limited capability in capturing hierarchical semantic relationships, and poor generalization across subjects. To address these issues, this study proposes a spatio-temporal multi-level adaptive cross-subject interactive attention network (ST-MSIAN). First, the spatiotemporal feature enhancemant module (SFEM) employs parallel temporal and spatial branches to jointly model temporal dynamics and channel dependencies of EEG signals. Multi-scale dilated convolutions with dilation rates of 1, 2, and 4 are adopted to capture both local and long-range temporal dependencies, while spatial convolutions and attention mechanisms enhance inter-channel representations. Subsequently, subject-specific tokens are introduced to explicitly characterize individual neural patterns and reduce the influence of inter-subject variability. Based on these representations, the multi-level interactive attention mechanism (MIAM) progressively extracts hierarchical semantic information through channel-level, subject-level, and global-level attention interactions. The resulting features are further processed by a six-layer Transformer encoder to model contextual dependencies. Finally, a KAN-based projection head with learnable activation functions is employed to establish a more effective nonlinear mapping between high-dimensional EEG features and the language embedding space of a pretrained BART decoder, enabling accurate text generation. Experiments are conducted on the publicly available Zurich cognitive language processing corpus (ZuCo), including both ZuCo v1.0 and ZuCo v2.0 datasets. The proposed ST-MSIAN achieves a BLEU-1 score of 44.10%, a BLEU-3 score of 16.21%, a ROUGE-F score of 37.62%, and a BERTScore-F score of 60.08%. Compared with the strongest competing method, the proposed approach improves BLEU-1, ROUGE-F, and BERTScore-F by 2.23, 4.87, and 4.66 percentage points, respectively. The proposed ST-MSIAN provides an effective and robust solution for cross-subject neural language decoding and offers promising potential for future applications in assistive communication systems and intelligent brain-computer interfaces.

  • WANG Shuo, CHEN Zhe, YIN Fuliang

    2026(4):1026-1040, DOI: 10.16337/j.1004-9037.2026.04.008

    Abstract:

    Non-line-of-sight sound source localization refers to the technique of achieving accurate sound source positioning in environments with obstacles and occlusions by analyzing acoustic propagation characteristics and applying traditional digital signal processing or deep learning approaches. This technology has a wide range of application prospects in robot navigation, smart home, security monitoring and other fields. This paper presents a comprehensive review of non-line-of-sight sound source localization methods. Building on the fundamental concepts of non-line-of-sight localization, it classifies and discusses existing approaches and research progress according to three main modeling frameworks: Multi-sensor localization models, geometrical acoustics models, and statistical acoustics models. Moreover, this paper summarizes various non-line-of-sight source localization methods, identifies current challenges, and presents promising directions for future research.

  • WEN Jinying, XIE Yuelei, LIU Xiangguo

    2026(4):1041-1057, DOI: 10.16337/j.1004-9037.2026.04.009

    Abstract:

    In order to solve the problem of identifying unknown emitters in an open electromagnetic environment, we propose an open-set identification method based on the convolutional neural network- Fourier analysis network (CNN-FAN) model. The proposed method initially introduces the network layer of Fourier analysis networks (FAN). The Fourier series properties are utilized to effectively extract the frequency components of the signal, followed by its integration with a convolutional neural network to formulate the CNN-FAN network model. Subsequently, the Softmax layer is substituted with an open-set identification model, facilitated by the utilization of OpenMax. Finally, a joint optimization strategy combining center loss and cross-entropy loss is adopted to optimize model performance. This optimization reduces the feature distance of individual radiation sources and narrows inter-class gaps, which improves the classification performance and enables OpenMax to distinguish between known and unknown emitter categories. The proposed method is validated and subjected to experimental analysis on the open-source WiSig dataset under varying degrees of openness. Experimental results demonstrate that the proposed method attains a recognition rate of 95% at an openness level of 0.057 and an open-set recognition rate of 84% at an openness level of 0.184, thereby outperforming other open-set recognition methods. In addition, the proposed framework provides a systematic solution for open-set radio frequency fingerprint identification by jointly considering discriminative feature learning and unknown-class rejection. The CNN module captures local temporal characteristics from the input radio frequency signals, while the FAN layer further enhances the representation capability by modeling periodic and frequency-domain variations. This complementary feature extraction mechanism enables the network to obtain more robust and separable emitter-specific representations. Moreover, by incorporating OpenMax into the decision stage, the model is no longer restricted to closed-set classification and can assign samples from unseen emitters to unknown categories according to their activation distribution. The combination of feature compactness optimization and open-set probability calibration improves both known-emitter identification and unknown-emitter rejection. These results indicate that the proposed CNN-FAN-based open-set identification method has strong applicability in realistic electromagnetic environments where unknown emitters may appear during deployment.

  • PU Yunwei, ZHAO Huijie, LI Wenkang, TIAN Chunjin

    2026(4):1058-1077, DOI: 10.16337/j.1004-9037.2026.04.010

    Abstract:

    Radar emitter signal sorting (RESS) plays a critical role on modern electronic warfare. To address the limitations of traditional sorting methods, such as low efficiency and limited accuracy in complex electromagnetic environments, as well as radar-communication spectrum sharing (RCSS) brought about by the development of 5G communication technologies, this paper proposes an innovative solution based on multi-task learning and multi-domain feature fusion. First, the proposed method constructs a multi-task learning (MTL) framework to jointly preprocess noisy radar-communication mixed signals, including signal cleansing, denoising, and time-frequency feature enhancement. Second, a feature fusion network is designed to extract multi-scale features from both the time-frequency image domain and the time-delay-Doppler domain. An improved iterative attention mechanism is employed to achieve cross-domain feature fusion, resulting in a high-dimensional feature representation with a strong discriminative power. Finally, a clustering model based on DeepCluster is used to accurately sort radar emitter signals. Experiments demonstrates that under low signal-to-noise ratio (SNR) condition of -6 dB, the proposed method effectively suppresses communication signal interference and achieves a sorting accuracy of 93.8%. Compared with existing approaches, the proposed solution demonstrates significant advantages in enhancing the robustness of signal preprocessing and maintaining high sorting accuracy, underscoring its strong potential for practical engineering applications.

  • SUN Weibai, NI Shuyan, SONG Xin, LEI Tuofeng, CHEN Quan

    2026(4):1078-1088, DOI: 10.16337/j.1004-9037.2026.04.011

    Abstract:

    To address the performance degradation of Luby transform (LT) codes under Gaussian channels, where traditional robust soliton distribution (RSD) suffers from high decoding complexity and excessive overhead due to noise accumulation and redundancy in high-degree symbols, this paper proposes a dynamic coupled truncated distribution (DCTD). The proposed distribution first dynamically truncates the RSD to suppress redundant high-degree symbols, thereby reducing computational complexity. It then introduces a pre-optimized fixed distribution to further sparsify the probability density in the medium-to-high degree region. To achieve an adaptive and smooth fusion of the two distributions, the Kullback-Leibler (KL) divergence is employed to quantitatively measure the statistical difference between the truncated RSD and the fixed distribution, based on which an adaptive switching point selection algorithm is designed. Furthermore, a nonlinear mixing strategy is developed, incorporating a sinusoidal modulation term to enhance the robustness of low-degree symbols and an exponential correction term to suppress redundant high-degree connections, ensuring a smooth transition and improved probability allocation across critical degree ranges. Simulation results demonstrate that at a code length of 1 000, DCTD reduces the average decoding overhead by 21.5% compared to RSD while maintaining the same decoding success rate. Additionally, the average degree is reduced by 29.90% to 46.49% across various code lengths, indicating a significant reduction in decoding complexity. These results confirm that DCTD achieves a better trade-off between decoding overhead, complexity and success rate, providing an effective degree distribution solution for LT codes in noise-limited scenarios.

  • CHANG Yuhan, LI Xiaorong, ZHAO Hongfeng

    2026(4):1089-1102, DOI: 10.16337/j.1004-9037.2026.04.012

    Abstract:

    Blind decoding is the core technique to recover the original data sequence without any prior information in non-cooperative communication scenarios. However, the absence of prior information poses severe challenges to blind decoding, making the integrated implementation of blind recognition of channel coding scheme and blind decoding a critical and urgent problem to be solved in this field. In current wireless communication systems, limited and fixed standard protocols are widely adopted, and specific protocols have a strong binding relationship with their corresponding channel coding schemes. Therefore, the coding scheme adopted by the communication system can be inferred through the blind recognition of the communication protocol. To address the above challenges, this paper proposes a novel hybrid deep learning network model, and constructs an integrated solution for blind recognition of channel coding and blind decoding. Numerical experimental results demonstrate that, under low signal-to-noise ratio regimes, the proposed model achieves higher recognition accuracy of channel coding and stronger anti-noise robustness compared with conventional schemes. This work provides a new technical approach for the research of blind decoding technology, and also supplies theoretical and experimental support for engineering applications in the field of information countermeasures.

  • ZHAO Xiaohong, CHEN Hao, SHI Guanju, GUAN Ti, DING Xuhui, ZHU Chao

    2026(4):1103-1117, DOI: 10.16337/j.1004-9037.2026.04.013

    Abstract:

    Generating realistic network traffic for digital twin systems is difficult because traffic packets exhibit temporal dependence, periodic variation, and scenario-driven uncertainty. To address this problem, this paper proposes FlowDiff, a diffusion-based packet generation algorithm, and TDGM, a temporal diffusion generation model. FlowDiff formulates packet generation as a conditioned reverse-diffusion process that progressively denoises latent traffic features. TDGM integrates temporal-aware embedding, periodic features, convolutional neural network (CNN) and Transformer modules, and cross-attention to capture local, global, and long-range traffic patterns. Experiments on real network traffic datasets show that the generated samples are close to real traffic in MSE, KL divergence, FID, and statistical features, demonstrating the effectiveness of the method for digital twin network simulation.

  • LIU Zhongkai, ZHANG Nan

    2026(4):1118-1132, DOI: 10.16337/j.1004-9037.2026.04.014

    Abstract:

    Attribute reduction is one of the core research problems in rough set theory, aiming to identify a more concise subset of attributes without compromising the classification capability of the original dataset. Among various reduction strategies, attribute reduction based on generalized decision preservation has been widely applied due to its strong interpretability and practical applicability. However, existing algorithms for generalized decision-preserving attribute reduction often suffer from significant computational inefficiency when dealing with large-scale datasets. How to efficiently achieve attribute reduction under the generalized decision-preserving framework has become a critical issue that demands urgent attention. To address this problem, this paper explores the stability characteristics of objects in the positive region during the generalized decision-preserving process. By employing the stripped quotient set method, a simplified and more representative equivalence class structure is constructed. On this basis, we propose an efficient heuristic attribute reduction algorithm under the generalized decision preservation framework. The proposed algorithm significantly improves computational efficiency while ensuring the correctness of the reduction results. Experimental results on eight benchmark datasets from University of California, Irvine (UCI) repository demonstrate that the proposed algorithm achieves higher reduction efficiency compared to existing methods.

  • ZHOU Ying, CHEN Yuming

    2026(4):1133-1146, DOI: 10.16337/j.1004-9037.2026.04.015

    Abstract:

    In few-shot image classification, the number of samples is usually small. Moreover, existing metric learning models cannot well balance local and global information. They are often affected by complex backgrounds and thus cannot focus well on the main subject information. To address these issues, a few-shot classification model named dual-granularity region confusion networks (DGRCN) that integrates region-reorganized images and sub-block features is proposed. The DGRCN model introduces the multi-granularity idea and proposes the region confusion granulation (RCG) method. Through this method, multi-granularity region-reorganized images and multi-granularity sub-block sequences are obtained, which alleviates the problem of insufficient samples. Secondly, picture-level and patch-level features are extracted from the region-reorganized images and sub-block sequences, respectively. An attention mechanism is added to refine the areas of focus and perform multi-granularity feature fusion based on these two types of features. Finally, the classification results of the dual features are adaptively fused, taking into account both global and local key information. DGRCN is experimented on the public datasets miniImageNet, CUB and Stanford Dogs. Compared with the baseline model DN4, the accuracy has been slightly improved. Experimental results show that DGRCN can effectively improve the performance of few-shot classification tasks.

  • LIU Shangdong, WU Jian, XU He, JI Yimu

    2026(4):1147-1163, DOI: 10.16337/j.1004-9037.2026.04.016

    Abstract:

    The conventional cycle generative adversarial network (CycleGAN), as an image style transfer model that does not require paired datasets, in the task of low-light domain image enhancement, there are problems such as the loss of details in the generated images, color distortion, and poor adaptability in complex scenarios. Based on this, this paper proposes an improved CycleGAN model which aims to enhance the effect of low-light image enhancement. First, in the design of the generator, a two-stage color correction module is integrated during the upsampling phase to alleviate the problems of color distortion and quality degradation in low-light environments. Second, the channel-spatial hybrid attention mechanism is embedded in the skip connection layer to achieve the adaptive strengthening of key information during the feature fusion process. Then, in the design of the discriminator, a global-local discrimination mode is adopted, enabling it to take into account the discrimination ability of both global information and local details. Finally, in the design of the loss function, perceptual style loss and content loss are introduced on the basis of adversarial loss to further improve the structural fidelity and visual naturalness of the generated images. Through the subjective and objective experimental evaluations and comparison with various representative image enhancement models, results show that this model has a good enhancement effect on low-light images. It can effectively enhance the overall brightness and local details of images without causing color distortion, thus improving the quality of the generated images.

  • ZHAO Shuai, LI Lin

    2026(4):1164-1177, DOI: 10.16337/j.1004-9037.2026.04.017

    Abstract:

    To capture the complex spatiotemporal dependencies in pedestrian trajectories, this paper proposes a trajectory prediction model that combines a bidirectional temporal learning module and a spatiotemporal interaction learning module. The model leverages bidirectional temporal feature modeling and self-supervised learning to extract spatiotemporal interaction features. In the bidirectional temporal learning module, a bidirectional temporal convolutional network is utilized to simultaneously model both historical and future trajectory information, enabling the capture of dynamic trajectory changes. In the spatiotemporal interaction learning module, the test-time training (TTT) layer is employed with a self-supervised learning mechanism to dynamically adjust feature representations during the inference stage, thereby modeling spatiotemporal correlations. Finally, an adaptive fusion strategy is used to combine the features extracted by the two modules, focusing on key features while suppressing irrelevant information. Experimental results demonstrate that the proposed model achieves competitive prediction performance on the ETH and UCY datasets.Highlights:1.Proposes a dual-branch framework fusing BiTCN bidirectional temporal modeling and TTT-based spatiotemporal self-supervised learning for multi-pedestrian trajectory prediction.2.BiTCN with bidirectional convolutions captures both historical and future trajectory dependencies, addressing incomplete feature representation from unidirectional temporal modeling.3.First introduces test-time training (TTT) layers to dynamically adapt feature representations via inference-time self-supervision, effectively modeling complex dynamic pedestrian interactions.4.Achieves state-of-the-art ADE/FDE performance on ETH and UCY benchmarks, outperforming mainstream baselines while maintaining lightweight design with only 22.4% parameters and 9% inference time of Social-LSTM.

  • XU Chudi, HE Hong, CHEN Jiayu, CHEN Yucong, LI Zexu

    2026(4):1178-1193, DOI: 10.16337/j.1004-9037.2026.04.018

    Abstract:

    Brain tumors are one of the most common diseases in the nervous system, and magnetic resonance imaging (MRI) examination is a common screening method for brain tumors. Although deep learning models have achieved high performance in MRI image recognition, their operational effectiveness is unsatisfactory in scenarios with limited computational resources, and their performance tends to decline after pruning for lightweight. Therefore, this study focuses on improving the brain tumor MRI image classification performance of lightweight models. To address the aforementioned challenges, classification method of variable-temperature self-distillation brain tumor MRI images based on label-auto-calibration is proposed. Starting from the core requirement of enhancing the performance of lightweight models, this study employs a loss function based on label-auto-calibration loss to avoid imparting incorrect knowledge from the self-teacher model during the self-distillation process. Additionally, a temperature-resetting variable-temperature distillation mechanism is introduced. When the validation accuracy does not improve in recent rounds of training, the temperature is reset based on the current training situation to prevent the lightweight network from encountering excessive learning difficulty. Experimental analysis on multiple brain tumor datasets demonstrates that the proposed self-distillation method can effectively enhance the brain tumor recognition performance of lightweight models, and its performance surpasses other typical self-distillation methods. This method requires no changes to the model structure, is easy to implement, and can be used to enhance the performance of lightweight networks.

  • JIN Kaiyu, HU Xiao, XIE Hong, LIAN Defu

    2026(4):1194-1211, DOI: 10.16337/j.1004-9037.2026.04.019

    Abstract:

    Remote sensing image-text retrieval has received increasing attention in recent years. Most existing methods rely on aligning global features between images and texts to perform retrieval, with a strong focus on feature extraction and fusion. However, they typically overlook the alignment between fine-grained object regions in images and semantic phrases in texts (i.e., local alignment). This limitation is particularly problematic in remote sensing, where intra-modality samples often exhibit high visual similarity, confusing global alignment. To address this, we propose GLAPA (Global and local alignment with phrase augmentation), a novel framework that complements global alignment with fine-grained local alignment. Specifically, GLAPA extracts conceptual phrases from textual descriptions and dynamically aligns them with relevant image patches, enhancing the model’s ability to capture detailed cross-modal semantics. In addition, we incorporate masked modeling to strengthen intra-modal feature learning, further improving retrieval performance. Extensive experiments on the RSICD and RSITMD datasets—overing ablation studies, hyperparameter analysis, and visualization—demonstrate that GLAPA significantly outperforms state-of-the-art methods. Visualization results also confirm the effectiveness of our local alignment strategy.

  • XU Yingzhe, DU Qingzhi, SHAO Yubin, DUO Lin

    2026(4):1212-1225, DOI: 10.16337/j.1004-9037.2026.04.020

    Abstract:

    In the field of traffic sign detection, challenges arise due to the small area coverage of distant traffic signs in the scene and the diverse scales of the signs. To overcome the above challenges, this paper presents an improved YOLOv8s-based traffic sign detection algorithm, YOLOv8s-REMN. First, the method introduces the receptive field attention convolution (RFAConv) into the backbone network to enhance the receptive field and feature extraction capability of the network. Second, the efficient attention-guided feature module(EAGFM) module is added to the neck network to optimize multi-scale feature fusion. Then, the multi-scale detail enhancement fusion (MSDEF) module is incorporated into the detection head to increase the small object detection head, improving the detection of small targets. Finally, the normalized Wasserstein distance (NWD) loss function replaces the CIoU loss function to optimize the bounding box regression and improve the precision of small object localization. Experimental results show that YOLOv8s-REMN achieves significant performance improvements on the TT100K dataset. Compared to the original YOLOv8s, mAP@0.5 increases by 6.6% and mAP@0.5:0.95 increases by 5.1%. The effectiveness of the algorithm is also validated on the Chinese Traffic Sign Detection dataset, CCTSDB2021, where YOLOv8s-REMN outperforms YOLOv8s with a 2.9% increase in mAP@0.5 and a 2.9% increase in mAP@0.5:0.95.

  • ZHANG Ling, ZHENG Ying, SONG Wanli

    2026(4):1226-1238, DOI: 10.16337/j.1004-9037.2026.04.021

    Abstract:

    Comment spams pose a significant threat to the reputation of e-commerce platforms. Existing methods based on graph neural networks for detecting spam comments face issues such as data sparsity, complex node relationships, and incomplete characterizations, which affect the detection performance. To address these problems, this paper proposes a graph neural network-based spam detection method, named GNNSD. The proposed method introduces a semantic-aware node enhancement mechanism to generate semantically relevant neighbors for minority class nodes to alleviate the performance decline caused by class imbalance and insufficient features. Further, GNNSD employs multi-head relation-aware attention and semantic-relation dual attention mechanisms to comprehensively capture the interaction patterns and fine-grained dependency features between nodes, and adopts cross-layer residual connection and contrastive learning strategies to enhance feature reuse and increase the discriminative power of node embeddings. Experiments are conducted on benchmark datasets, and the results show that GNNSD outperforms existing mainstream methods in terms of AUC, Micro-F1, and Macro-F1 metrics, achieves better detection performance in the case of class imbalance, and has a good potential in the field of spam detection.

WeChat
Quick search
Search term
Search word
From To
Volume retrieval