{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-channel-attention-selection-gan-with","title":"Multi-Channel Attention Selection GAN with Cascaded Semantic Guidance for Cross-View Image Translation","arxiv_id":"1904.06807","date":"2019-04-15","proceeding":"CVPR 2019 6","authors":["Hao Tang","Dan Xu","Nicu Sebe","Yanzhi Wang","Jason J. Corso","Yan Yan"],"abstract":"Cross-view image translation is challenging because it involves images with\ndrastically different views and severe deformation. In this paper, we propose a\nnovel approach named Multi-Channel Attention SelectionGAN (SelectionGAN) that\nmakes it possible to generate images of natural scenes in arbitrary viewpoints,\nbased on an image of the scene and a novel semantic map. The proposed\nSelectionGAN explicitly utilizes the semantic information and consists of two\nstages. In the first stage, the condition image and the target semantic map are\nfed into a cycled semantic-guided generation network to produce initial coarse\nresults. In the second stage, we refine the initial results by using a\nmulti-channel attention selection mechanism. Moreover, uncertainty maps\nautomatically learned from attentions are used to guide the pixel loss for\nbetter network optimization. Extensive experiments on Dayton, CVUSA and Ego2Top\ndatasets show that our model is able to generate significantly better results\nthan the state-of-the-art methods. The source code, data and trained models are\navailable at https://github.com/Ha0Tang/SelectionGAN.","url_abs":"http://arxiv.org/abs/1904.06807v2","url_pdf":"http://arxiv.org/pdf/1904.06807v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-channel-attention-selection-gan-with","repo_url":"https://github.com/Ha0Tang/HandGestureRecognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"multi-channel-attention-selection-gan-with","repo_url":"https://github.com/Ha0Tang/LocalGlobalGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"multi-channel-attention-selection-gan-with","repo_url":"https://github.com/Ha0Tang/SelectionGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"bird-view-synthesis","task_name":"Bird View Synthesis"},{"task_slug":"cross-view-image-to-image-translation","task_name":"Cross-View Image-to-Image Translation"},{"task_slug":"image-to-image-translation","task_name":"Image-to-Image Translation"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-view-image-to-image-translation-on-2","task":"Cross-View Image-to-Image Translation","dataset":"Dayton (256×256) - aerial-to-ground","model":"SelectionGAN","rank_in_archive_order":1,"of":6,"metrics":{"SSIM":"0.5938"},"uses_additional_data":false},{"leaderboard":"/sota/cross-view-image-to-image-translation-on-3","task":"Cross-View Image-to-Image Translation","dataset":"Dayton (256×256) - ground-to-aerial","model":"SelectionGAN","rank_in_archive_order":1,"of":4,"metrics":{"SSIM":"0.3284"},"uses_additional_data":false},{"leaderboard":"/sota/cross-view-image-to-image-translation-on-1","task":"Cross-View Image-to-Image Translation","dataset":"Dayton (64x64) - ground-to-aerial","model":"SelectionGAN","rank_in_archive_order":1,"of":5,"metrics":{"SSIM":"0.5118"},"uses_additional_data":false},{"leaderboard":"/sota/cross-view-image-to-image-translation-on","task":"Cross-View Image-to-Image Translation","dataset":"Dayton (64×64) - aerial-to-ground","model":"SelectionGAN","rank_in_archive_order":1,"of":5,"metrics":{"SSIM":"0.6865"},"uses_additional_data":false},{"leaderboard":"/sota/cross-view-image-to-image-translation-on-5","task":"Cross-View Image-to-Image Translation","dataset":"Ego2Top","model":"SelectionGAN","rank_in_archive_order":1,"of":4,"metrics":{"SSIM":"0.6024"},"uses_additional_data":false},{"leaderboard":"/sota/cross-view-image-to-image-translation-on-4","task":"Cross-View Image-to-Image Translation","dataset":"cvusa","model":"SelectionGAN","rank_in_archive_order":2,"of":7,"metrics":{"SSIM":"0.5323"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.06807","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}