{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-generalized-zero-shot-learners-for","title":"Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization","arxiv_id":"2302.00275","date":"2023-02-01","proceeding":null,"authors":["Lukas Haas","Silas Alberti","Michal Skreta"],"abstract":"Image geolocalization is the challenging task of predicting the geographic coordinates of origin for a given photo. It is an unsolved problem relying on the ability to combine visual clues with general knowledge about the world to make accurate predictions across geographies. We present $\\href{https://huggingface.co/geolocal/StreetCLIP}{\\text{StreetCLIP}}$, a robust, publicly available foundation model not only achieving state-of-the-art performance on multiple open-domain image geolocalization benchmarks but also doing so in a zero-shot setting, outperforming supervised models trained on more than 4 million images. Our method introduces a meta-learning approach for generalized zero-shot learning by pretraining CLIP from synthetic captions, grounding CLIP in a domain of choice. We show that our method effectively transfers CLIP's generalized zero-shot capabilities to the domain of image geolocalization, improving in-domain generalized zero-shot performance without finetuning StreetCLIP on a fixed set of classes.","url_abs":"https://arxiv.org/abs/2302.00275v1","url_pdf":"https://arxiv.org/pdf/2302.00275v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-generalized-zero-shot-learners-for","repo_url":"https://huggingface.co/geolocal/StreetCLIP","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"generalized-zero-shot-learning","task_name":"Generalized Zero-Shot Learning"},{"task_slug":"meta-learning","task_name":"Meta-Learning"},{"task_slug":"photo-geolocation-estimation","task_name":"Photo geolocation estimation"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"}],"methods":[{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/photo-geolocation-estimation-on-im2gps","task":"Photo geolocation estimation","dataset":"Im2GPS","model":"StreetCLIP (Zero-Shot)","rank_in_archive_order":11,"of":11,"metrics":{"City level (25 km)":"28.3","Continent level (2500 km)":"88.2","Country level (750 km)":"74.7","Reference images":"0","Region level (200 km)":"45.1","Training images":"1.1M"},"uses_additional_data":false},{"leaderboard":"/sota/photo-geolocation-estimation-on-im2gps3k","task":"Photo geolocation estimation","dataset":"Im2GPS3k","model":"StreetCLIP (Zero-Shot)","rank_in_archive_order":14,"of":14,"metrics":{"City level (25 km)":" 22.4","Continent level (2500 km)":" 80.4","Country level (750 km)":"61.3","Region level (200 km)":" 37.4","Street level (1 km)":"-","Training Images":"1.1M"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2302.00275","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}