{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/effovpr-effective-foundation-model","title":"EffoVPR: Effective Foundation Model Utilization for Visual Place Recognition","arxiv_id":"2405.18065","date":"2024-05-28","proceeding":null,"authors":["Issar Tzachor","Boaz Lerner","Matan Levy","Michael Green","Tal Berkovitz Shalev","Gavriel Habib","Dvir Samuel","Noam Korngut Zailer","Or Shimshi","Nir Darshan","Rami Ben-Ari"],"abstract":"The task of Visual Place Recognition (VPR) is to predict the location of a query image from a database of geo-tagged images. Recent studies in VPR have highlighted the significant advantage of employing pre-trained foundation models like DINOv2 for the VPR task. However, these models are often deemed inadequate for VPR without further fine-tuning on VPR-specific data. In this paper, we present an effective approach to harness the potential of a foundation model for VPR. We show that features extracted from self-attention layers can act as a powerful re-ranker for VPR, even in a zero-shot setting. Our method not only outperforms previous zero-shot approaches but also introduces results competitive with several supervised methods. We then show that a single-stage approach utilizing internal ViT layers for pooling can produce global features that achieve state-of-the-art performance, with impressive feature compactness down to 128D. Moreover, integrating our local foundation features for re-ranking further widens this performance gap. Our method also demonstrates exceptional robustness and generalization, setting new state-of-the-art performance, while handling challenging conditions such as occlusion, day-night transitions, and seasonal variations.","url_abs":"https://arxiv.org/abs/2405.18065v2","url_pdf":"https://arxiv.org/pdf/2405.18065v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"re-ranking","task_name":"Re-Ranking"},{"task_slug":"visual-place-recognition","task_name":"Visual Place Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-place-recognition-on-amstertime","task":"Visual Place Recognition","dataset":"AmsterTime","model":"EffoVPR","rank_in_archive_order":2,"of":8,"metrics":{"Recall@1":"65.5"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-eynsham","task":"Visual Place Recognition","dataset":"Eynsham","model":"EffoVPR","rank_in_archive_order":6,"of":7,"metrics":{"Recall@1":"91.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-mapillary-test","task":"Visual Place Recognition","dataset":"Mapillary test","model":"EffoVPR","rank_in_archive_order":5,"of":12,"metrics":{"Recall@1":"79.0","Recall@10":"91.6","Recall@5":"89.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-mapillary-val","task":"Visual Place Recognition","dataset":"Mapillary val","model":"EffoVPR","rank_in_archive_order":7,"of":18,"metrics":{"Recall@1":"92.8","Recall@10":"97.4","Recall@5":"97.2"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-nordland","task":"Visual Place Recognition","dataset":"Nordland","model":"EffoVPR","rank_in_archive_order":2,"of":13,"metrics":{"Recall@1":"95.0","Recall@5":"98.6"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-pittsburgh-30k","task":"Visual Place Recognition","dataset":"Pittsburgh-30k-test","model":"EffoVPR","rank_in_archive_order":5,"of":22,"metrics":{"Recall@1":"93.9","Recall@5":"97.4"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-sf-xl-test-v1","task":"Visual Place Recognition","dataset":"SF-XL test v1","model":"EffoVPR","rank_in_archive_order":1,"of":5,"metrics":{"Recall@1":"95.5","Recall@10":"98.1"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-sf-xl-test-v2","task":"Visual Place Recognition","dataset":"SF-XL test v2","model":"EffoVPR","rank_in_archive_order":2,"of":5,"metrics":{"Recall@1":"94.5","Recall@10":"97.8","Recall@5":"98.2"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-san-francisco","task":"Visual Place Recognition","dataset":"San Francisco Landmark Dataset","model":"EffoVPR","rank_in_archive_order":2,"of":3,"metrics":{"Recall@1":"93.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-st-lucia","task":"Visual Place Recognition","dataset":"St Lucia","model":"EffoVPR","rank_in_archive_order":1,"of":14,"metrics":{"Recall@1":"100.0","Recall@5":"100.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-place-recognition-on-tokyo247","task":"Visual Place Recognition","dataset":"Tokyo247","model":"EffoVPR","rank_in_archive_order":2,"of":14,"metrics":{"Recall@1":"98.7","Recall@10":"98.7","Recall@5":"98.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2405.18065","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}