{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/georanker-distance-aware-ranking-for","title":"GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization","arxiv_id":"2505.13731","date":"2025-05-19","proceeding":null,"authors":["Pengyue Jia","Seongheon Park","Song Gao","Xiangyu Zhao","Yixuan Li"],"abstract":"Worldwide image geolocalization-the task of predicting GPS coordinates from images taken anywhere on Earth-poses a fundamental challenge due to the vast diversity in visual content across regions. While recent approaches adopt a two-stage pipeline of retrieving candidates and selecting the best match, they typically rely on simplistic similarity heuristics and point-wise supervision, failing to model spatial relationships among candidates. In this paper, we propose GeoRanker, a distance-aware ranking framework that leverages large vision-language models to jointly encode query-candidate interactions and predict geographic proximity. In addition, we introduce a multi-order distance loss that ranks both absolute and relative distances, enabling the model to reason over structured spatial relationships. To support this, we curate GeoRanking, the first dataset explicitly designed for geographic ranking tasks with multimodal candidate information. GeoRanker achieves state-of-the-art results on two well-established benchmarks (IM2GPS3K and YFCC4K), significantly outperforming current best methods.","url_abs":"https://arxiv.org/abs/2505.13731v1","url_pdf":"https://arxiv.org/pdf/2505.13731v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"photo-geolocation-estimation","task_name":"Photo geolocation estimation"}],"methods":[{"method_slug":"adopt","method_name":"ADOPT"},{"method_slug":"gps","method_name":"GPS"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/photo-geolocation-estimation-on-im2gps3k","task":"Photo geolocation estimation","dataset":"Im2GPS3k","model":"GeoRanker","rank_in_archive_order":1,"of":14,"metrics":{"City level (25 km)":"45.05","Continent level (2500 km)":"89.29","Country level (750 km)":"76.31","Region level (200 km)":"61.49","Street level (1 km)":"18.79","Training Images":"100K"},"uses_additional_data":false},{"leaderboard":"/sota/photo-geolocation-estimation-on-yfcc4k","task":"Photo geolocation estimation","dataset":"YFCC4k","model":"GeoRanker","rank_in_archive_order":2,"of":4,"metrics":{"City (25 km)":"43.54","Continent (2500 km)":"82.45","Country (750 km)":"69.79","Median Error (km)":"/","Region (200 km)":"54.32","Street (1 km)":"32.94"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2505.13731","atlas_url":"https://app.syntology.ai/?focus=2505.13731","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}