Papers › Global–Local Information Soft-Alignment for Cross-Modal Remote-Sensing Image–Text Retrieval

Global–Local Information Soft-Alignment for Cross-Modal Remote-Sensing Image–Text Retrieval

14 May 2024journal 2024 5archive 2025-07-28

Gang Hu, Zaidao Wen, Yafei Lv, Jianting Zhang, Qian Wu

Cross-modal remote-sensing image–text retrieval (CMRSITR) is a challenging task that aims to retrieve target remote-sensing (RS) images based on textual descriptions. However, the modal gap between texts and RS images poses a significant challenge. RS images comprise multiple targets and complex backgrounds, necessitating the mining of both global and local information (GaLR) for effective CMRSITR. Existing approaches primarily focus on local image features while disregarding the local features of the text and their correspondence. These methods typically fuse global and local image features and align them with global text features. However, they struggle to eliminate the influence of cluttered backgrounds and may overlook crucial targets. To address these limitations, we propose a novel framework for CMRSITR based on a transformer architecture, which leverages global–local information soft alignment (GLISA) to enhance retrieval performance. Our framework incorporates a global image extraction module, which captures the global semantic features of image–text pairs and effectively represents the relationships among multiple targets in RS images. In addition, we introduce an adaptive local information extraction (ALIE) module that adaptively mines discriminative local clues from both RS images and texts, aligning the corresponding fine-grained information. To mitigate semantic ambiguities during the alignment of local features, we design a local information soft-alignment (LISA) module. In comparative evaluations using two public CMRSITR datasets, our proposed method achieves state-of-the-art results, surpassing not only traditional cross-modal retrieval methods by a substantial margin but also other contrastive language-image pretraining (CLIP)-based methods.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Modal RetrievalCross-Modal Retrieval on RSITMDImage-text RetrievalRetrievalText Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval RSICD GLISA Image-to-text R@1 20.68% #3 of 10 Archive leaderboard report
Cross-Modal Retrieval RSICD GLISA Mean Recall 37.69% #3 of 10 Archive leaderboard report
Cross-Modal Retrieval RSICD GLISA text-to-image R@1 14.73% #3 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD GLISA Image-to-text R@1 32.08% #3 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD GLISA Mean Recall 50.69% #3 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD GLISA text-to-imageR@1 23.36% #3 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ALIGNFocus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections