Papers › Learning Cross-Scale Visual Representations for Real-Time Image Geo-Localization

Learning Cross-Scale Visual Representations for Real-Time Image Geo-Localization

9 Sep 2021arXiv:2109.04087archive 2025-07-28

Tianyi Zhang, Matthew Johnson-Roberson

Robot localization remains a challenging task in GPS denied environments. State estimation approaches based on local sensors, e.g. cameras or IMUs, are drifting-prone for long-range missions as error accumulates. In this study, we aim to address this problem by localizing image observations in a 2D multi-modal geospatial map. We introduce the cross-scale dataset and a methodology to produce additional data from cross-modality sources. We propose a framework that learns cross-scale visual representations without supervision. Experiments are conducted on data from two different domains, underwater and aerial. In contrast to existing studies in cross-view image geo-localization, our approach a) performs better on smaller-scale multi-modal maps; b) is more computationally efficient for real-time applications; c) can serve directly in concert with state estimation pipelines.

PaperPDFCode

Code

tyz1030/croscalerep officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

State Estimationgeo-localization

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

GPS

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections