{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/localizing-and-orienting-street-views-using","title":"Localizing and Orienting Street Views Using Overhead Imagery","arxiv_id":"1608.00161","date":"2016-07-30","proceeding":null,"authors":["Nam Vo","James Hays"],"abstract":"In this paper we aim to determine the location and orientation of a\nground-level query image by matching to a reference database of overhead (e.g.\nsatellite) images. For this task we collect a new dataset with one million\npairs of street view and overhead images sampled from eleven U.S. cities. We\nexplore several deep CNN architectures for cross-domain matching --\nClassification, Hybrid, Siamese, and Triplet networks. Classification and\nHybrid architectures are accurate but slow since they allow only partial\nfeature precomputation. We propose a new loss function which significantly\nimproves the accuracy of Siamese and Triplet embedding networks while\nmaintaining their applicability to large-scale retrieval tasks like image\ngeolocalization. This image matching task is challenging not just because of\nthe dramatic viewpoint difference between ground-level and overhead imagery but\nbecause the orientation (i.e. azimuth) of the street views is unknown making\ncorrespondence even more difficult. We examine several mechanisms to match in\nspite of this -- training for rotation invariance, sampling possible rotations\nat query time, and explicitly predicting relative rotation of ground and\noverhead images with our deep networks. It turns out that explicit orientation\nsupervision also improves location prediction accuracy. Our best performing\narchitectures are roughly 2.5 times as accurate as the commonly used Siamese\nnetwork baseline.","url_abs":"http://arxiv.org/abs/1608.00161v2","url_pdf":"http://arxiv.org/pdf/1608.00161v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":null,"task_name":"Triplet"}],"methods":[],"datasets_introduced":[{"slug":"dayton","name":"Dayton","full_name":"Dayton"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1608.00161","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}