Papers › Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

21 Nov 2017CVPR 2018 6arXiv:1711.07613archive 2025-07-28

Qi Wu, Peng Wang, Chunhua Shen, Ian Reid, Anton Van Den Hengel

The Visual Dialogue task requires an agent to engage in a conversation about an image with a human. It represents an extension of the Visual Question Answering task in that the agent needs to answer a question about an image, but it needs to do so in light of the previous dialogue that has taken place. The key challenge in Visual Dialogue is thus maintaining a consistent, and natural dialogue while continuing to answer questions correctly. We present a novel approach that combines Reinforcement Learning and Generative Adversarial Networks (GANs) to generate more human-like responses to questions. The GAN helps overcome the relative paucity of training data, and the tendency of the typical MLE-based approach to generate overly terse answers. Critically, the GAN is tightly integrated into the attention mechanism that generates human-interpretable reasons for each answer. This means that the discriminative model of the GAN has the task of assessing whether a candidate answer is generated by a human or not, given the provided reason. This is significant because it drives the generative model to produce high quality answers that are well supported by the associated reasoning. The method also generates the state-of-the-art results on the primary benchmark.

PaperPDFConference PDF

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringReinforcement LearningVisual DialogVisual Question AnsweringVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Dialog VisDial v0.9 val CoAtt MRR 63.98 #4 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val CoAtt Mean Rank 4.47 #4 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val CoAtt R@1 50.29 #4 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val CoAtt R@10 88.81 #4 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val CoAtt R@5 80.71 #4 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Convolution

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections