Papers › LatteGAN: Visually Guided Language Attention for Multi-Turn Text-Conditioned Image Manipulation

LatteGAN: Visually Guided Language Attention for Multi-Turn Text-Conditioned Image Manipulation

28 Dec 2021arXiv:2112.13985archive 2025-07-28

Shoya Matsumori, Yuki Abe, Kosuke Shingyouchi, Komei Sugiura, Michita Imai

Text-guided image manipulation tasks have recently gained attention in the vision-and-language community. While most of the prior studies focused on single-turn manipulation, our goal in this paper is to address the more challenging multi-turn image manipulation (MTIM) task. Previous models for this task successfully generate images iteratively, given a sequence of instructions and a previously generated image. However, this approach suffers from under-generation and a lack of generated quality of the objects that are described in the instructions, which consequently degrades the overall performance. To overcome these problems, we present a novel architecture called a Visually Guided Language Attention GAN (LatteGAN). Here, we address the limitations of the previous approaches by introducing a Visually Guided Language Attention (Latte) module, which extracts fine-grained text representations for the generator, and a Text-Conditioned U-Net discriminator architecture, which discriminates both the global and local representations of fake or real images. Extensive experiments on two distinct MTIM datasets, CoDraw and i-CLEVR, demonstrate the state-of-the-art performance of the proposed model.

PaperPDFCode

Code

smatsumori/lattegan officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image ManipulationText-to-Image Generation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text-to-Image Generation GeNeVA (CoDraw) LatteGAN F1-score 77.51± 0.52 #1 of 2 Archive leaderboard report
Text-to-Image Generation GeNeVA (CoDraw) LatteGAN rsim 54.16± 0.21 #1 of 2 Archive leaderboard report
Text-to-Image Generation GeNeVA (i-CLEVR) LatteGAN F1-score 97.26±1.56 #1 of 2 Archive leaderboard report
Text-to-Image Generation GeNeVA (i-CLEVR) LatteGAN rsim 83.21± 1.70 #1 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Concatenated Skip ConnectionConvolutionMax PoolingReLUU-Net

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections