Papers › Multi-Modal Open-Domain Dialogue

Multi-Modal Open-Domain Dialogue

2 Oct 2020EMNLP 2021 11arXiv:2010.01082archive 2025-07-28

Kurt Shuster, Eric Michael Smith, Da Ju, Jason Weston

Recent work in open-domain conversational agents has demonstrated that significant improvements in model engagingness and humanness metrics can be achieved via massive scaling in both pre-training data and model size (Adiwardana et al., 2020; Roller et al., 2020). However, if we want to build agents with human-like abilities, we must expand beyond handling just text. A particularly important topic is the ability to see images and communicate about what is perceived. With the goal of engaging humans in multi-modal dialogue, we investigate combining components from state-of-the-art open-domain dialogue agents with those from state-of-the-art vision models. We study incorporating different image fusion schemes and domain-adaptive pre-training and fine-tuning strategies, and show that our best resulting model outperforms strong existing models in multi-modal dialogue while simultaneously performing as well as its predecessor (text-only) BlenderBot (Roller et al., 2020) in text-based conversation. We additionally investigate and incorporate safety components in our final model, and show that such efforts do not diminish model performance with respect to engagingness metrics.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Visual Dialog

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Dialog BlendedSkillTalk Multi-Modal BlenderBot BLEU-4 1 #1 of 1 Archive leaderboard report
Visual Dialog BlendedSkillTalk Multi-Modal BlenderBot F1 17.8 #1 of 1 Archive leaderboard report
Visual Dialog BlendedSkillTalk Multi-Modal BlenderBot ROUGE-L 19.3 #1 of 1 Archive leaderboard report
Visual Dialog ConvAI2 Multi-Modal BlenderBot BLEU-4 1.1 #1 of 1 Archive leaderboard report
Visual Dialog ConvAI2 Multi-Modal BlenderBot F1 18.4 #1 of 1 Archive leaderboard report
Visual Dialog ConvAI2 Multi-Modal BlenderBot ROUGE-L 22.6 #1 of 1 Archive leaderboard report
Visual Dialog EmpatheticDialogues Multi-Modal BlenderBot BLEU-4 1.5 #1 of 1 Archive leaderboard report
Visual Dialog EmpatheticDialogues Multi-Modal BlenderBot F1 19.2 #1 of 1 Archive leaderboard report
Visual Dialog EmpatheticDialogues Multi-Modal BlenderBot ROUGE-L 24.5 #1 of 1 Archive leaderboard report
Visual Dialog Image-Chat Multi-Modal BlenderBot BLEU-4 40 #1 of 1 Archive leaderboard report
Visual Dialog Image-Chat Multi-Modal BlenderBot F1 13.1 #1 of 1 Archive leaderboard report
Visual Dialog Image-Chat Multi-Modal BlenderBot ROUGE-L 18 #1 of 1 Archive leaderboard report
Visual Dialog Wizard of Wikipedia Multi-Modal BlenderBot BLEU-4 2.2 #1 of 1 Archive leaderboard report
Visual Dialog Wizard of Wikipedia Multi-Modal BlenderBot F1 18.6 #1 of 1 Archive leaderboard report
Visual Dialog Wizard of Wikipedia Multi-Modal BlenderBot ROUGE-L 17.4 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections