{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lattegan-visually-guided-language-attention","title":"LatteGAN: Visually Guided Language Attention for Multi-Turn Text-Conditioned Image Manipulation","arxiv_id":"2112.13985","date":"2021-12-28","proceeding":null,"authors":["Shoya Matsumori","Yuki Abe","Kosuke Shingyouchi","Komei Sugiura","Michita Imai"],"abstract":"Text-guided image manipulation tasks have recently gained attention in the vision-and-language community. While most of the prior studies focused on single-turn manipulation, our goal in this paper is to address the more challenging multi-turn image manipulation (MTIM) task. Previous models for this task successfully generate images iteratively, given a sequence of instructions and a previously generated image. However, this approach suffers from under-generation and a lack of generated quality of the objects that are described in the instructions, which consequently degrades the overall performance. To overcome these problems, we present a novel architecture called a Visually Guided Language Attention GAN (LatteGAN). Here, we address the limitations of the previous approaches by introducing a Visually Guided Language Attention (Latte) module, which extracts fine-grained text representations for the generator, and a Text-Conditioned U-Net discriminator architecture, which discriminates both the global and local representations of fake or real images. Extensive experiments on two distinct MTIM datasets, CoDraw and i-CLEVR, demonstrate the state-of-the-art performance of the proposed model.","url_abs":"https://arxiv.org/abs/2112.13985v2","url_pdf":"https://arxiv.org/pdf/2112.13985v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lattegan-visually-guided-language-attention","repo_url":"https://github.com/smatsumori/lattegan","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-manipulation","task_name":"Image Manipulation"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"}],"methods":[{"method_slug":"concatenated-skip-connection","method_name":"Concatenated Skip Connection"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"u-net","method_name":"U-Net"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-to-image-generation-on-geneva-codraw","task":"Text-to-Image Generation","dataset":"GeNeVA (CoDraw)","model":"LatteGAN","rank_in_archive_order":1,"of":2,"metrics":{"F1-score":"77.51± 0.52","rsim":"54.16± 0.21"},"uses_additional_data":false},{"leaderboard":"/sota/text-to-image-generation-on-geneva-i-clevr","task":"Text-to-Image Generation","dataset":"GeNeVA (i-CLEVR)","model":"LatteGAN","rank_in_archive_order":1,"of":2,"metrics":{"F1-score":"97.26±1.56","rsim":"83.21± 1.70"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}