Papers › Learning to Embed Multi-Modal Contexts for Situated Conversational Agents
Learning to Embed Multi-Modal Contexts for Situated Conversational Agents
Haeju Lee, Oh Joon Kwon, Yunseon Choi, Minho Park, Ran Han, Yoonhyung Kim, Jinhyeon Kim, Youngjune Lee, Haebin Shin, Kangwook Lee, Kee-Eung Kim
The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of four subtasks, multi-modal disambiguation (MM-Disamb), multi-modal coreference resolution (MM-Coref), multi-modal dialog state tracking (MM-DST), and response retrieval and generation. While many task-oriented dialog systems usually tackle each subtask separately, we propose a jointly learned multi-modal encoder-decoder that incorporates visual inputs and performs all four subtasks at once for efficiency. This approach won the MM-Coref and response retrieval subtasks and nominated runner-up for the remaining subtasks using a single unified model at the 10th Dialog Systems Technology Challenge (DSTC10), setting a high bar for the novel task of multi-modal task-oriented dialog systems.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Dialogue State Tracking | SIMMC2.0 | BART-large | Act F1 | 96.3 | #2 of 5 | Archive leaderboard | report |
| Dialogue State Tracking | SIMMC2.0 | BART-large | Slot F1 | 88.3 | #2 of 5 | Archive leaderboard | report |
| Dialogue State Tracking | SIMMC2.0 | BART-base | Act F1 | 95.2 | #3 of 5 | Archive leaderboard | report |
| Dialogue State Tracking | SIMMC2.0 | BART-base | Slot F1 | 82.0 | #3 of 5 | Archive leaderboard | report |
| Response Generation | SIMMC2.0 | BART-large | BLEU | 33.1 | #2 of 5 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections