{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-ground-level-scene-layout-from","title":"Predicting Ground-Level Scene Layout from Aerial Imagery","arxiv_id":"1612.02709","date":"2016-12-08","proceeding":"CVPR 2017 7","authors":["Menghua Zhai","Zachary Bessinger","Scott Workman","Nathan Jacobs"],"abstract":"We introduce a novel strategy for learning to extract semantically meaningful\nfeatures from aerial imagery. Instead of manually labeling the aerial imagery,\nwe propose to predict (noisy) semantic features automatically extracted from\nco-located ground imagery. Our network architecture takes an aerial image as\ninput, extracts features using a convolutional neural network, and then applies\nan adaptive transformation to map these features into the ground-level\nperspective. We use an end-to-end learning approach to minimize the difference\nbetween the semantic segmentation extracted directly from the ground image and\nthe semantic segmentation predicted solely based on the aerial image. We show\nthat a model learned using this strategy, with no additional training, is\nalready capable of rough semantic labeling of aerial imagery. Furthermore, we\ndemonstrate that by finetuning this model we can achieve more accurate semantic\nsegmentation than two baseline initialization strategies. We use our network to\naddress the task of estimating the geolocation and geoorientation of a ground\nimage. Finally, we show how features extracted from an aerial image can be used\nto hallucinate a plausible ground-level panorama.","url_abs":"http://arxiv.org/abs/1612.02709v1","url_pdf":"http://arxiv.org/pdf/1612.02709v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"predicting-ground-level-scene-layout-from","repo_url":"https://github.com/viibridges/crossnet","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"cross-view-image-to-image-translation","task_name":"Cross-View Image-to-Image Translation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-view-image-to-image-translation-on-4","task":"Cross-View Image-to-Image Translation","dataset":"cvusa","model":"CrossNet","rank_in_archive_order":6,"of":7,"metrics":{"SSIM":"0.4147"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1612.02709","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}