Papers › Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms

Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms

25 Sep 2024arXiv:2409.16850archive 2025-07-28

Chun-Jung Lin, Sourav Garg, Tat-Jun Chin, Feras Dayoub

We present a novel method for scene change detection that leverages the robust feature extraction capabilities of a visual foundational model, DINOv2, and integrates full-image cross-attention to address key challenges such as varying lighting, seasonal variations, and viewpoint differences. In order to effectively learn correspondences and mis-correspondences between an image pair for the change detection task, we propose to a) ``freeze'' the backbone in order to retain the generality of dense foundation features, and b) employ ``full-image'' cross-attention to better tackle the viewpoint variations between the image pair. We evaluate our approach on two benchmark datasets, VL-CMU-CD and PSCD, along with their viewpoint-varied versions. Our experiments demonstrate significant improvements in F1-score, particularly in scenarios involving geometric changes between image pairs. The results indicate our method's superior generalization capabilities over existing state-of-the-art approaches, showing robustness against photometric and geometric variations as well as better overall generalization when fine-tuned to adapt to new environments. Detailed ablation studies further validate the contributions of each component in our architecture. Source code will be made publicly available upon acceptance.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

ChadLin9596/Robust-Scene-Change-Detection officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Change DetectionScene Change Detection

Datasets

Introduced by this paper, per the archive.

Unaligned-VL-CMU-CD (neighbor distance 2)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Scene Change Detection Unaligned-VL-CMU-CD (neighbor distance 2) Robust-Scene-Change-Detection (Diff-View Augmentation) F1-score 0.784 #1 of 2 Archive leaderboard report
Scene Change Detection Unaligned-VL-CMU-CD (neighbor distance 2) Robust-Scene-Change-Detection F1-score 0.739 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections