Abstract:Existing point cloud semantic segmentation methods for railway scenes struggle to simultaneously balance cross-scenario robustness and object segmentation accuracy. Using RandLA-Net as the basic framework, this paper develops an implicit-explicit 3D structure adaptive fusion network (RandLA-Net with Adaptive Fusion of Explicit Structure, RandLA-AFES) specifically tailored for actual railway engineering scenarios. In the encoder stage, an ex-plicit structural feature extraction module is introduced in parallel, which utilizes a local spatial partitioning strategy to capture structural kernels and generate explicit geometric priors, effectively compensating for the fine-grained structural loss during the downsampling process. Meanwhile, a neighborhood feature sampling module is utilized to aggregate deep local contextual information, ensuring the integrity of the high-dimensional implicit semantic rep-resentation of the point cloud. Furthermore, a dynamic adaptive fusion mechanism is designed. By introducing learnable weight parameters, it dynamically calibrates the fusion ratio of explicit geometric features and implicit semantic features, achieving the optimal coupling of dual-modal features and effectively suppressing redundant noise. Extensive experiments on the public railway scene dataset WHU-Railway3D demonstrate that, compared with the baseline model, the proposed method outperforms existing mainstream methods in metrics such as mean intersection over union (mIoU) and overall accuracy (OA) with only a minimal increase in parameters and inference time. Moreover, it exhibits excellent semantic segmentation capabilities across various complex scenarios, verifying its effectiveness and potential for engineering applications.