Abstract:Remote Sensing Image Change Captioning (RSICC) aims to generate natural language descriptions that reflect land-surface alterations using multi-temporal images, holding substantial value in domains such as land-resource monitoring and disaster assessment. However, in practical deployment, bi-temporal images are frequently vulnerable to complex non-semantic interference, such as seasonal successions and sudden illumination fluctuations, which easily induces "semantic hallucinations" in deep networks. Moreover, the ubiquitous simple differencing mechanisms adopted by existing methods struggle to effectively suppress these confounding factors and lack sufficient sensitivity to highly localized change signals, thereby yielding imprecise descriptive captions with poor contextual alignment. To tackle these multi-faceted challenges, this paper proposes a robust remote sensing image change captioning framework grounded on multi-scale feature compensation and change enhancement. To effectively alleviate the severe interference stemming from pseudo-change factors like lighting variations and environmental dynamics, a novel Feature Compensation Network (FCN) is designed. Leveraging a bidirectional guided compensation mechanism via cross-temporal attention, the FCN adaptively modulates bi-temporal features to enforce strict semantic and stylistic consistency across non-change background regions, thereby achieving the explicit suppression of illumination and shadow pseudo-changes. Concurrently, to overcome the perceptual deficiency regarding subtle change signals, a Discriminative Change-Aware Guided Module (DAGM) and a multi-scale fusion strategy are introduced. By synchronously reinforcing highly discriminative representations across both spatial and channel dimensions, this module substantially elevates the framework's operational sensitivity to the minor and intricate evolution of complex ground objects. Comprehensive quantitative comparison experiments conducted on the challenging LEVIR-CC and Dubai-CC benchmark datasets demonstrate that the proposed method consistently and significantly outperforms mainstream state-of-the-art algorithms across multiple core evaluation metrics, such as BLEU-4 and CIDEr, validating its distinct superiority in generating accurate and grammatically coherent text descriptions. Furthermore, qualitative analysis and comprehensive visualization results provide intuitive, empirical evidence supporting the robust cross-temporal interpretability of the model. Meanwhile, rigorous ablation studies systematically verify the theoretical necessity and practical effectiveness of each key proposed component. Ultimately, both experimental and theoretical investigations indicate that the proposed Feature Compensation Network (FCN) successfully achieves bidirectional interactive alignment and explicit interference suppression within non-change regions. Meanwhile, the seamlessly integrated Change-Aware Guided Module (DAGM) and multi-scale feature fusion strategy can effectively mine, preserve, and reinforce fine-grained details within genuine change regions, offering a highly promising and generalizable solution for high-precision earth observation and semantic change interpretation tasks.