{"url":"/sota/question-answering-on-coqa","task":{"name":"Question Answering","url":"/task/question-answering","note":null},"dataset":{"name":"CoQA","url":"/dataset/coqa"},"category":"Natural Language Processing","categories":["Miscellaneous","Natural Language Processing","Reasoning"],"category_note":null,"description":"Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include [SQuAD](/dataset/squad), [HotPotQA](/dataset/hotpotqa), [bAbI](/dataset/babi-1), [TriviaQA](/dataset/triviaqa), [WikiQA](/dataset/wikiqa), and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.\r\n\r\n( Image credit: [SQuAD](https://rajpurkar.github.io/mlx/qa-and-squad/) )","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["In-domain","Out-of-domain","Overall"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"In-domain":null,"Out-of-domain":null,"Overall":null}},"counts":{"rows":9,"rows_with_code":9,"rows_with_paper_page":9,"rows_dated":9,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"BERT Large Augmented (single model)","metrics":{"In-domain":"82.5","Out-of-domain":"77.6","Overall":"81.1"},"uses_additional_data":false,"paper_date":"2018-10-11","paper":"/paper/bert-pre-training-of-deep-bidirectional","paper_url":"https://arxiv.org/abs/1810.04805v2","paper_title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","code":"https://github.com/huggingface/transformers","n_code_links":534,"syntology":{"n_ran":204,"n_unverified":455,"n_samples":659,"n_pointer_only_licence":149}},{"rank_in_archive_order":2,"model":"BERT-base finetune (single model)","metrics":{"In-domain":"79.8","Out-of-domain":"74.1","Overall":"78.1"},"uses_additional_data":false,"paper_date":"2018-10-11","paper":"/paper/bert-pre-training-of-deep-bidirectional","paper_url":"https://arxiv.org/abs/1810.04805v2","paper_title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","code":"https://github.com/huggingface/transformers","n_code_links":534,"syntology":{"n_ran":204,"n_unverified":455,"n_samples":659,"n_pointer_only_licence":149}},{"rank_in_archive_order":3,"model":"BiDAF++ (single model)","metrics":{"In-domain":"69.4","Out-of-domain":"63.8","Overall":"67.8"},"uses_additional_data":false,"paper_date":"2018-09-27","paper":"/paper/a-qualitative-comparison-of-coqa-squad-20-and","paper_url":"https://arxiv.org/abs/1809.10735v2","paper_title":"A Qualitative Comparison of CoQA, SQuAD 2.0 and QuAC","code":"https://github.com/my89/co-squac","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"DrQA + seq2seq with copy attention (single model)","metrics":{"In-domain":"67.0","Out-of-domain":"60.4","Overall":"65.1"},"uses_additional_data":false,"paper_date":"2018-08-21","paper":"/paper/coqa-a-conversational-question-answering","paper_url":"http://arxiv.org/abs/1808.07042v2","paper_title":"CoQA: A Conversational Question Answering Challenge","code":"https://github.com/stanfordnlp/coqa-baselines","n_code_links":4,"syntology":{"n_ran":2,"n_unverified":0,"n_samples":2,"n_pointer_only_licence":0}},{"rank_in_archive_order":5,"model":"Vanilla DrQA (single model)","metrics":{"In-domain":"54.5","Out-of-domain":"47.9","Overall":"52.6"},"uses_additional_data":false,"paper_date":"2018-08-21","paper":"/paper/coqa-a-conversational-question-answering","paper_url":"http://arxiv.org/abs/1808.07042v2","paper_title":"CoQA: A Conversational Question Answering Challenge","code":"https://github.com/stanfordnlp/coqa-baselines","n_code_links":4,"syntology":{"n_ran":2,"n_unverified":0,"n_samples":2,"n_pointer_only_licence":0}},{"rank_in_archive_order":6,"model":"FlowQA (single model)","metrics":{"Out-of-domain":"71.8","Overall":"75.0"},"uses_additional_data":false,"paper_date":"2018-10-06","paper":"/paper/flowqa-grasping-flow-in-history-for","paper_url":"http://arxiv.org/abs/1810.06683v3","paper_title":"FlowQA: Grasping Flow in History for Conversational Machine Comprehension","code":"https://github.com/momohuang/FlowQA","n_code_links":1,"syntology":null},{"rank_in_archive_order":7,"model":"GPT-3 175B (few-shot, k=32)","metrics":{"Overall":"85"},"uses_additional_data":false,"paper_date":"2020-05-28","paper":"/paper/language-models-are-few-shot-learners","paper_url":"https://arxiv.org/abs/2005.14165v4","paper_title":"Language Models are Few-Shot Learners","code":"https://github.com/ggml-org/llama.cpp","n_code_links":67,"syntology":{"n_ran":15,"n_unverified":50,"n_samples":65,"n_pointer_only_licence":4}},{"rank_in_archive_order":8,"model":"SDNet (ensemble)","metrics":{"Overall":"79.3"},"uses_additional_data":false,"paper_date":"2018-12-10","paper":"/paper/sdnet-contextualized-attention-based-deep","paper_url":"http://arxiv.org/abs/1812.03593v5","paper_title":"SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering","code":"https://github.com/Microsoft/SDNet","n_code_links":6,"syntology":null},{"rank_in_archive_order":9,"model":"SDNet (single model)","metrics":{"Overall":"76.6"},"uses_additional_data":false,"paper_date":"2018-12-10","paper":"/paper/sdnet-contextualized-attention-based-deep","paper_url":"http://arxiv.org/abs/1812.03593v5","paper_title":"SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering","code":"https://github.com/Microsoft/SDNet","n_code_links":6,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":5,"rows_with_any_sample_ran":5,"distinct_papers_with_graph_line":3,"distinct_papers_with_any_sample_ran":3,"samples_over_distinct_papers":{"n_ran":221,"n_unverified":505,"n_samples":726,"n_pointer_only_licence":153,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":427,"n_unverified":960,"n_samples":1387,"n_pointer_only_licence":302,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}