基于短语增强的全局与局部对齐的遥感图像文本检索方法
作者:
作者单位:

认知智能全国重点实验室(中国科学技术大学计算机科学与技术学院), 合肥 230088

作者简介:

通讯作者:

基金项目:

国家科技创新2030重大项目(2021ZD0111800)。


Global and Local Alignment with Phrase Augmentation for Remote Sensing Image-Text Retrieval
Author:
Affiliation:

Laboratory of Cognitive Intelligence (School of Computer Science and Technology, University of Science and Technology of China), Hefei 230088, China

Fund Project:

National Key R&D Program of China (No.2021ZD0111800).

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
    摘要:

    遥感图文检索现有的方法主要通过对齐配对图像和文本的全局特征来实现检索,但大多方法的注意力都集中在特征提取和特征融合方面,在特征对齐时都只关注于经过汇聚后的全局特征之间的对齐(全局对齐),而忽略掉了图像中细粒度的物体特征与文本中的语义短语特征之间的对齐(局部对齐)。然而在遥感领域,同一个模态内的样本往往高度相似,这会干扰全局对齐。本文提出一种全新的模型——带有短语增强的全局与局部对齐(Global and local alignment with phrase augmentation,GLAPA),在对齐全局特征的同时,引入细粒度的局部特征对齐作为补充。本文通过提取文本中的概念短语,并动态地与图像中的区域特征进行局部对齐,显著提升了检索性能。此外,本文还通过掩码建模增强了单模态内特征之间的关联性学习,进一步提升遥感图像-文本检索的效果。本文在RSICD和RSITMD数据集上进行了大量实验,包括消融实验、超参数敏感性分析和可视化分析等。实验结果表明,GLAPA在RSICD和RSITMD数据集上的表现优于现有的方法,可视化实验也验证了局部对齐的准确性。

    Abstract:

    Remote sensing image-text retrieval has received increasing attention in recent years. Most existing methods rely on aligning global features between images and texts to perform retrieval, with a strong focus on feature extraction and fusion. However, they typically overlook the alignment between fine-grained object regions in images and semantic phrases in texts (i.e., local alignment). This limitation is particularly problematic in remote sensing, where intra-modality samples often exhibit high visual similarity, confusing global alignment. To address this, we propose GLAPA (Global and local alignment with phrase augmentation), a novel framework that complements global alignment with fine-grained local alignment. Specifically, GLAPA extracts conceptual phrases from textual descriptions and dynamically aligns them with relevant image patches, enhancing the model’s ability to capture detailed cross-modal semantics. In addition, we incorporate masked modeling to strengthen intra-modal feature learning, further improving retrieval performance. Extensive experiments on the RSICD and RSITMD datasets—overing ablation studies, hyperparameter analysis, and visualization—demonstrate that GLAPA significantly outperforms state-of-the-art methods. Visualization results also confirm the effectiveness of our local alignment strategy.

    参考文献
    相似文献
    引证文献
引用本文

金开宇,胡筱,谢洪,连德富.基于短语增强的全局与局部对齐的遥感图像文本检索方法[J].数据采集与处理,2026,(4):1194-1211

复制
分享
文章指标
  • 点击次数:
  • 下载次数:
历史
  • 收稿日期:2025-03-01
  • 最后修改日期:2025-09-08
  • 录用日期:
  • 在线发布日期: 2026-08-13