Papers › GSV-Cities: Toward Appropriate Supervised Visual Place Recognition
GSV-Cities: Toward Appropriate Supervised Visual Place Recognition
Amar Ali-bey, Brahim Chaib-Draa, Philippe Giguère
This paper aims to investigate representation learning for large scale visual place recognition, which consists of determining the location depicted in a query image by referring to a database of reference images. This is a challenging task due to the large-scale environmental changes that can occur over time (i.e., weather, illumination, season, traffic, occlusion). Progress is currently challenged by the lack of large databases with accurate ground truth. To address this challenge, we introduce GSV-Cities, a new image dataset providing the widest geographic coverage to date with highly accurate ground truth, covering more than 40 cities across all continents over a 14-year period. We subsequently explore the full potential of recent advances in deep metric learning to train networks specifically for place recognition, and evaluate how different loss functions influence performance. In addition, we show that performance of existing methods substantially improves when trained on GSV-Cities. Finally, we introduce a new fully convolutional aggregation layer that outperforms existing techniques, including GeM, NetVLAD and CosPlace, and establish a new state-of-the-art on large-scale benchmarks, such as Pittsburgh, Mapillary-SLS, SPED and Nordland. The dataset and code are available for research purposes at https://github.com/amaralibey/gsv-cities.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Place Recognition | Mapillary val | Conv-AP | Recall@1 | 83.4 | #13 of 18 | Archive leaderboard | report |
| Visual Place Recognition | Mapillary val | Conv-AP | Recall@10 | 92.3 | #13 of 18 | Archive leaderboard | report |
| Visual Place Recognition | Mapillary val | Conv-AP | Recall@5 | 90.5 | #13 of 18 | Archive leaderboard | report |
| Visual Place Recognition | Nordland | Conv-AP | Recall@1 | 38.5 | #12 of 13 | Archive leaderboard | report |
| Visual Place Recognition | Nordland | Conv-AP | Recall@5 | 53.9 | #12 of 13 | Archive leaderboard | report |
| Visual Place Recognition | Pittsburgh-250k-test | Conv-AP | Recall@1 | 92.4 | #10 of 13 | Archive leaderboard | report |
| Visual Place Recognition | Pittsburgh-250k-test | Conv-AP | Recall@10 | 98.6 | #10 of 13 | Archive leaderboard | report |
| Visual Place Recognition | Pittsburgh-250k-test | Conv-AP | Recall@5 | 97.6 | #10 of 13 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections