Papers › Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing...

Reducing Semantic Confusion: Scene-aware Aggregation Network for Remote Sensing Cross-modal Retrieval

12 Jun 2023ICMR 2023 6archive 2025-07-28

Jiancheng Pan, Qing Ma, Cong Bai

Recently, remote sensing cross-modal retrieval has received incredible attention from researchers. However, the unique nature of remote-sensing images leads to many semantic confusion zones in the semantic space, which greatly affects retrieval performance. We propose a novel scene-aware aggregation network (SWAN) to reduce semantic confusion by improving scene perception capability. In visual representation, a visual multiscale fusion module (VMSF) is presented to fuse visual features with different scales as a visual representation backbone. Meanwhile, a scene fine-grained sensing module (SFGS) is proposed to establish the associations of salient features at different granularity. A scene-aware visual aggregation representation is formed by the visual information generated by these two modules. In textual representation, a textual coarse-grained enhancement module (TCGE) is designed to enhance the semantics of text and to align visual information. Furthermore, as the diversity and differentiation of remote sensing scenes weaken the understanding of scenes, a new metric, namely, scene recall is proposed to measure the perception of scenes by evaluating scene-level retrieval performance, which can also verify the effectiveness of our approach in reducing semantic confusion. By performance comparisons, ablation studies and visualization analysis, we validated the effectiveness and superiority of our approach on two datasets, RSICD and RSITMD. The source code is available at https://github.com/kinshingpoon/SWAN-pytorch.

PaperPDFCode

Code

kinshingpoon/swan-pytorch officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Modal RetrievalRetrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval RSICD SWAN Image-to-text R@1 7.41% #8 of 10 Archive leaderboard report
Cross-Modal Retrieval RSICD SWAN Mean Recall 20.61% #8 of 10 Archive leaderboard report
Cross-Modal Retrieval RSICD SWAN text-to-image R@1 5.56% #8 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD SWAN Image-to-text R@1 13.35% #9 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD SWAN Mean Recall 34.11% #9 of 10 Archive leaderboard report
Cross-Modal Retrieval RSITMD SWAN text-to-imageR@1 11.24% #9 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ALIGN

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections