Papers › A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval

A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval

27 Oct 2023ACMMM 2023 10archive 2025-07-28

Jiancheng Pan, Qing Ma, Cong Bai

This paper presents a prior instruction representation framework (PIR) for remote sensing image-text retrieval, aimed at remote sensing vision-language understanding tasks to solve the semantic noise problem. Our highlight is the proposal of a paradigm that draws on prior knowledge to instruct adaptive learning of vision and text representations. Concretely, two progressive attention encoder (PAE) structures, Spatial-PAE and Temporal-PAE, are proposed to perform long-range dependency modeling to enhance key feature representation. In vision representation, Vision Instruction Representation (VIR) based on Spatial-PAE exploits the prior-guided knowledge of the remote sensing scene recognition by building a belief matrix to select key features for reducing the impact of semantic noise. In text representation, Language Cycle Attention (LCA) based on Temporal-PAE uses the previous time step to cyclically activate the current time step to enhance text representation capability. A cluster-wise affiliation loss is proposed to constrain the inter-classes and to reduce the semantic confusion zones in the common subspace. Comprehensive experiments demonstrate that using prior knowledge instruction could enhance vision and text representations and could outperform the state-of-the-art methods on two benchmark datasets, RSICD and RSITMD.

PaperPDFCode

Code

jaychempan/PIR officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Modal RetrievalImage-text RetrievalRetrievalScene RecognitionText Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval RSICD PIR Image-to-text R@1 9.88% #6 of 10 Archive leaderboard report
Cross-Modal Retrieval RSICD PIR Mean Recall 24.46% #6 of 10 Archive leaderboard report
Cross-Modal Retrieval RSICD PIR text-to-image R@1 6.97% #6 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD PIR Image-to-text R@1 18.14% #6 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD PIR Mean Recall 38.24% #6 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD PIR text-to-imageR@1 12.17% #6 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections