Papers › DoQA -- Accessing Domain-Specific FAQs via Conversational QA

DoQA -- Accessing Domain-Specific FAQs via Conversational QA

4 May 2020arXiv:2005.01328archive 2025-07-28

Jon Ander Campos, Arantxa Otegi, Aitor Soroa, Jan Deriu, Mark Cieliebak, Eneko Agirre

The goal of this work is to build conversational Question Answering (QA) interfaces for the large body of domain-specific information available in FAQ sites. We present DoQA, a dataset with 2,437 dialogues and 10,917 QA pairs. The dialogues are collected from three Stack Exchange sites using the Wizard of Oz method with crowdsourcing. Compared to previous work, DoQA comprises well-defined information needs, leading to more coherent and natural conversations with less factoid questions and is multi-domain. In addition, we introduce a more realistic information retrieval(IR) scenario where the system needs to find the answer in any of the FAQ documents. The results of an existing, strong, system show that, thanks to transfer learning from a Wikipedia QA dataset and fine tuning on a single FAQ domain, it is possible to build high quality conversational QA systems for FAQs without in-domain training data. The good results carry over into the more challenging IR scenario. In both cases, there is still ample room for improvement, as indicated by the higher human upperbound.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Conversational Question AnsweringInformation RetrievalQuestion AnsweringRetrievalTransfer Learning

Datasets

Introduced by this paper, per the archive.

DoQA

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Wizard

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections