Datasets › CoQA

CoQA (Conversational Question Answering Challenge)

Introduced by Siva Reddy et al. in CoQA: A Conversational Question Answering Challenge1 Jan 2018 archive 2025-07-28

CoQA is a large-scale dataset for building Conversational Question Answering systems. The goal of the CoQA challenge is to measure the ability of machines to understand a text passage and answer a series of interconnected questions that appear in a conversation.

CoQA contains 127,000+ questions with answers collected from 8000+ conversations. Each conversation is collected by pairing two crowdworkers to chat about a passage in the form of questions and answers. The unique features of CoQA include 1) the questions are conversational; 2) the answers can be free-form text; 3) each answer also comes with an evidence subsequence highlighted in the passage; and 4) the passages are collected from seven diverse domains. CoQA has a lot of challenging phenomena not present in existing reading comprehension datasets, e.g., coreference and pragmatic reasoning.

Source: https://stanfordnlp.github.io/coqa/ Image Source: https://stanfordnlp.github.io/coqa/

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering CoQA BERT Large Augmented (single model) In-domain 82.5 BERT: Pre-training of Deep Bidirectional Transformers... huggingface/transformers +533 9 Compare
Generative Question Answering CoQA ERNIE-GEN F1-Score 84.5 ERNIE-GEN: An Enhanced Multi-Flow Pre-training and... PaddlePaddle/PaddleNLP +4 3 Compare

Papers archive 2025-07-28

8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 281. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Language Models are Few-Shot Learners 67 1 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning Framework for Natural Language Generation 5 1 26 Jan 2020 not harvested
Unified Language Model Pre-training for Natural Language Understanding and Generation 9 1 8 May 2019 not harvested
SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering 6 2 10 Dec 2018 not harvested
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 534 2 11 Oct 2018 ran 204 of 659 samples (455 unverified; 149 pointer-only for licence)
FlowQA: Grasping Flow in History for Conversational Machine Comprehension 1 1 6 Oct 2018 not harvested
A Qualitative Comparison of CoQA, SQuAD 2.0 and QuAC 1 1 27 Sep 2018 not harvested
CoQA: A Conversational Question Answering Challenge 4 3 21 Aug 2018 ran 2 of 2 samples (0 unverified)

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (multiple)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CoQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections