Browse State-of-the-Art › Visual Place Recognition

Visual Place Recognition

141 papers with code · 40 benchmarks · 28 datasets archive 2025-07-28

Computer Vision

Visual Place Recognition is the task of matching a view of a place with a different view of the same place taken at a different time.

Source: Visual place recognition using landmark distribution descriptors

Image credit: Visual place recognition using landmark distribution descriptors

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

40 leaderboard tables shown for this task, 40 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 40 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
Pittsburgh-30k-test (22 rows) Pair-VPR-p Pair-VPR: Place-Aware Pre-training and Contrastive Pair... code — Compare
Mapillary val (18 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
St Lucia (14 rows) EffoVPR EffoVPR: Effective Foundation Model Utilization for Visual Place... — — Compare
Tokyo247 (14 rows) Pair-VPR-p Pair-VPR: Place-Aware Pre-training and Contrastive Pair... code — Compare
Nordland (13 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
Pittsburgh-250k-test (13 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
Mapillary test (12 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
17 Places (8 rows) SegVLAD-FineT (M) Revisit Anything: Visual Place Recognition via Image Segment Retrieval code Syntology ran 6 of 9 samples · 3 unverified Compare
AmsterTime (8 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
Baidu Mall (8 rows) SegVLAD-PreT (M) Revisit Anything: Visual Place Recognition via Image Segment Retrieval code Syntology ran 6 of 9 samples · 3 unverified Compare
Nardo-Air R (8 rows) AnyLoc-VLAD-DINO AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
Eynsham (7 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
Gardens Point (7 rows) AnyLoc-VLAD-DINOv2 AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
Hawkins (7 rows) AnyLoc-VLAD-DINOv2 AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
Laurel Caverns (7 rows) AnyLoc-VLAD-DINOv2 AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
Mid-Atlantic Ridge (7 rows) AnyLoc-VLAD-DINOv2 AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
Nardo-Air (7 rows) AnyLoc-VLAD-DINOv2 AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
Oxford RobotCar Dataset (7 rows) AnyLoc-VLAD-DINOv2 AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
SPED (7 rows) BoQ BoQ: A Place is Worth a Bag of Learnable Queries code Syntology ran 2 of 2 samples · 0 unverified Compare
VP-Air (7 rows) AnyLoc-VLAD-DINOv2 AnyLoc: Towards Universal Visual Place Recognition code Syntology ran 1 of 1 samples · 0 unverified Compare
KITTI360pose (5 rows) MambaPlace MambaPlace:Text-to-Point-Cloud Cross-Modal Place Recognition with... code — Compare
SF-XL test v1 (5 rows) EffoVPR EffoVPR: Effective Foundation Model Utilization for Visual Place... — — Compare
SF-XL test v2 (5 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
Berlin Kudamm (4 rows) HEAPUtil A Hierarchical Dual Model of Environment- and Place-Specific... code — Compare
Nordland* (2760 queries) (4 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
SVOX-Night (4 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
SVOX-Overcast (4 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
SVOX-Rain (4 rows) QAA-DINOv2-B-8192 Query-Based Adaptive Aggregation for Multi-Dataset Joint Training... — — Compare
SVOX-Snow (4 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
SVOX-Sun (4 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
CV-Cities (3 rows) CV-Cities CV-Cities: Advancing Cross-View Geo-Localization in Global Cities code Syntology ran 0 of 9 samples · 9 unverified Compare
MSLS (3 rows) ProGEO ProGEO: Generating Prompts through Image-Text Contrastive Learning... code — Compare
Oxford RobotCar (LiDAR 4096 points+RGB) (3 rows) MinkLoc++ (LiDAR+RGB) MinkLoc++: Lidar and Monocular Image Fusion for Place Recognition code Syntology ran 2 of 5 samples · 3 unverified Compare
San Francisco Landmark Dataset (3 rows) BoQ BoQ: A Place is Worth a Bag of Learnable Queries code Syntology ran 2 of 2 samples · 0 unverified Compare
SF-XL Night (2 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
SF-XL Occlusion (2 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
SVOX (2 rows) FoL Focus on Local: Finding Reliable Discriminative Regions for Visual... code Syntology ran 3 of 3 samples · 0 unverified Compare
Inside Out (1 row) SegVLAD-FineT (M) Revisit Anything: Visual Place Recognition via Image Segment Retrieval code Syntology ran 6 of 9 samples · 3 unverified Compare
KITTI (1 row) None SSC: Semantic Scan Context for Large-Scale Place Recognition code — Compare
VP Air (1 row) SegVLAD-PreT (M) Revisit Anything: Visual Place Recognition via Image Segment Retrieval code Syntology ran 6 of 9 samples · 3 unverified Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

28 datasets whose archive record lists this task, ordered by the archive's paper count.

Subtasks archive 2025-07-28

3 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 141 papers with code (297 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections