Papers › DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue

DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue

17 Nov 2019arXiv:1911.07251archive 2025-07-28

Xiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang, Xingxing Zhang, Yue Hu, Qi Wu

Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any objects, relationships or semantics. The key challenge in Visual Dialogue task is thus to learn a more comprehensive and semantic-rich image representation which may have adaptive attentions on the image for variant questions. In this research, we propose a novel model to depict an image from both visual and semantic perspectives. Specifically, the visual view helps capture the appearance-level information, including objects and their relationships, while the semantic view enables the agent to understand high-level visual semantics from the whole image to the local regions. Futhermore, on top of such multi-view image features, we propose a feature selection framework which is able to adaptively capture question-relevant information hierarchically in fine-grained level. The proposed method achieved state-of-the-art results on benchmark Visual Dialogue datasets. More importantly, we can tell which modality (visual or semantic) has more contribution in answering the current question by visualizing the gate values. It gives us insights in understanding of human cognition in Visual Dialogue.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

JXZe/DualVD officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringVisual DialogVisual Question AnsweringVisual Question Answering (VQA)feature selection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Dialog VisDial v0.9 val DualVD MRR 62.94 #6 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val DualVD Mean Rank 4.17 #6 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val DualVD R@1 48.64 #6 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val DualVD R@10 89.94 #6 of 18 Archive leaderboard report
Visual Dialog VisDial v0.9 val DualVD R@5 80.89 #6 of 18 Archive leaderboard report
Visual Dialog Visual Dialog v1.0 test-std DualVD MRR (x 100) 63.23 #62 of 80 Archive leaderboard report
Visual Dialog Visual Dialog v1.0 test-std DualVD Mean 4.11 #62 of 80 Archive leaderboard report
Visual Dialog Visual Dialog v1.0 test-std DualVD NDCG (x 100) 56.32 #62 of 80 Archive leaderboard report
Visual Dialog Visual Dialog v1.0 test-std DualVD R@1 49.25 #62 of 80 Archive leaderboard report
Visual Dialog Visual Dialog v1.0 test-std DualVD R@10 89.7 #62 of 80 Archive leaderboard report
Visual Dialog Visual Dialog v1.0 test-std DualVD R@5 80.23 #62 of 80 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Feature Selection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections