Papers › Global–Local Information Soft-Alignment for Cross-Modal Remote-Sensing Image–Text Retrieval
Global–Local Information Soft-Alignment for Cross-Modal Remote-Sensing Image–Text Retrieval
Gang Hu, Zaidao Wen, Yafei Lv, Jianting Zhang, Qian Wu
Cross-modal remote-sensing image–text retrieval (CMRSITR) is a challenging task that aims to retrieve target remote-sensing (RS) images based on textual descriptions. However, the modal gap between texts and RS images poses a significant challenge. RS images comprise multiple targets and complex backgrounds, necessitating the mining of both global and local information (GaLR) for effective CMRSITR. Existing approaches primarily focus on local image features while disregarding the local features of the text and their correspondence. These methods typically fuse global and local image features and align them with global text features. However, they struggle to eliminate the influence of cluttered backgrounds and may overlook crucial targets. To address these limitations, we propose a novel framework for CMRSITR based on a transformer architecture, which leverages global–local information soft alignment (GLISA) to enhance retrieval performance. Our framework incorporates a global image extraction module, which captures the global semantic features of image–text pairs and effectively represents the relationships among multiple targets in RS images. In addition, we introduce an adaptive local information extraction (ALIE) module that adaptively mines discriminative local clues from both RS images and texts, aligning the corresponding fine-grained information. To mitigate semantic ambiguities during the alignment of local features, we design a local information soft-alignment (LISA) module. In comparative evaluations using two public CMRSITR datasets, our proposed method achieves state-of-the-art results, surpassing not only traditional cross-modal retrieval methods by a substantial margin but also other contrastive language-image pretraining (CLIP)-based methods.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Cross-Modal Retrieval | RSICD | GLISA | Image-to-text R@1 | 20.68% | #3 of 10 | Archive leaderboard | report |
| Cross-Modal Retrieval | RSICD | GLISA | Mean Recall | 37.69% | #3 of 10 | Archive leaderboard | report |
| Cross-Modal Retrieval | RSICD | GLISA | text-to-image R@1 | 14.73% | #3 of 10 | Archive leaderboard | report |
| Cross-Modal Retrieval | RSITMD | GLISA | Image-to-text R@1 | 32.08% | #3 of 10 | Archive leaderboard | report |
| Cross-Modal Retrieval | RSITMD | GLISA | Mean Recall | 50.69% | #3 of 10 | Archive leaderboard | report |
| Cross-Modal Retrieval | RSITMD | GLISA | text-to-imageR@1 | 23.36% | #3 of 10 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections