Papers › Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic...

Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction

24 Nov 2018ICCV 2019 10arXiv:1811.09845archive 2025-07-28

Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, Devon Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, Graham W. Taylor

Conditional text-to-image generation is an active area of research, with many possible applications. Existing research has primarily focused on generating a single image from available conditioning information in one step. One practical extension beyond one-step generation is a system that generates an image iteratively, conditioned on ongoing linguistic input or feedback. This is significantly more challenging than one-step generation tasks, as such a system must understand the contents of its generated images with respect to the feedback history, the current feedback, as well as the interactions among concepts present in the feedback history. In this work, we present a recurrent image generation model which takes into account both the generated output up to the current step as well as all past instructions for generation. We show that our model is able to generate the background, add new objects, and apply simple transformations to existing objects. We believe our approach is an important step toward interactive generation. Code and data is available at: https://www.microsoft.com/en-us/research/project/generative-neural-visual-artist-geneva/ .

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Maluuba/GeNeVA mentioned on GitHubpytorchNOASSERTION report
Maluuba/GeNeVA_datasets mentioned on GitHubNOASSERTION report
capstonecs42/GeNeVA_datasets_dev mentioned on GitHubNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Text-to-Image Generation GeNeVA (CoDraw) GeNeVA-GAN F1-score 58.83 #2 of 2 Archive leaderboard report
Text-to-Image Generation GeNeVA (CoDraw) GeNeVA-GAN rsim 35.41 #2 of 2 Archive leaderboard report
Text-to-Image Generation GeNeVA (i-CLEVR) GeNeVA-GAN F1-score 88.39 #2 of 2 Archive leaderboard report
Text-to-Image Generation GeNeVA (i-CLEVR) GeNeVA-GAN rsim 74.02 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections