Papers › Fusing task-oriented and open-domain dialogues in conversational agents

Fusing task-oriented and open-domain dialogues in conversational agents

9 Sep 2021arXiv:2109.04137archive 2025-07-28

Tom Young, Frank Xing, Vlad Pandelea, Jinjie Ni, Erik Cambria

The goal of building intelligent dialogue systems has largely been separately pursued under two paradigms: task-oriented dialogue (TOD) systems, which perform goal-oriented functions, and open-domain dialogue (ODD) systems, which focus on non-goal-oriented chitchat. The two dialogue modes can potentially be intertwined together seamlessly in the same conversation, as easily done by a friendly human assistant. Such ability is desirable in conversational agents, as the integration makes them more accessible and useful. Our paper addresses this problem of fusing TODs and ODDs in multi-turn dialogues. Based on the popular TOD dataset MultiWOZ, we build a new dataset FusedChat, by rewriting the existing TOD turns and adding new ODD turns. This procedure constructs conversation sessions containing exchanges from both dialogue modes. It features inter-mode contextual dependency, i.e., the dialogue turns from the two modes depend on each other. Rich dependency patterns including co-reference and ellipsis are features. The new dataset, with 60k new human-written ODD turns and 5k re-written TOD turns, offers a benchmark to test a dialogue model's ability to perform inter-mode conversations. This is a more challenging task since the model has to determine the appropriate dialogue mode and generate the response based on the inter-mode context. But such models would better mimic human-level conversation capabilities. We evaluate baseline models on this task, including classification-based two-stage models and two-in-one fused models. We publicly release FusedChat and the baselines to propel future work on inter-mode dialogue systems https://github.com/tomyoung903/FusedChat.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

tomyoung903/fusedchat officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Dialogue Generation

Datasets

Introduced by this paper, per the archive.

FusedChat

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Dialogue Generation FusedChat Classification-based model BLEU 12.17 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Inform 75.1 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Inform_mct 90.8 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Joint SA 0.600 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model PPL 10.50 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model SSA 0.55 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Sensibleness 0.58 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Slot Accuracy 0.973 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Specificity 0.51 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Success 60.9 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Classification-based model Success_mct 74.4 #1 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model BLEU 12.05 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Inform 70.4 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Inform_mct 90.1 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Joint SA 0.592 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model PPL 10.49 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model SSA 0.50 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Sensibleness 0.52 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Slot Accuracy 0.972 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Specificity 0.47 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Success 57.0 #2 of 2 Archive leaderboard report
Dialogue Generation FusedChat Two-in-one model Success_mct 72.7 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections