{"url":"/dataset/ambignq","name":"AmbigNQ","full_name":null,"description_markdown":"The **AmbigNQ dataset** is a valuable resource for exploring ambiguity in open-domain question answering. Let me provide you with some details:\r\n\r\n1. **Task Description**:\r\n   - Ambiguity is inherent in open-domain question answering, especially when dealing with new topics. It can be challenging to formulate questions that have a single, unambiguous answer.\r\n   - The **AmbigQA task** involves predicting a set of question-answer pairs, where each plausible answer is paired with a disambiguated rewrite of the original question.\r\n\r\n2. **Dataset Construction**:\r\n   - To study this task, the researchers constructed the **AmbigNQ dataset**.\r\n   - AmbigNQ covers **14,042 questions** from **NQ-open**, which is an existing open-domain QA benchmark.\r\n   - Surprisingly, over **half** of the questions in NQ-open exhibit ambiguity.\r\n   - The types of ambiguity are diverse and sometimes subtle, often becoming apparent only after examining evidence provided by a very large text corpus.\r\n\r\n3. **Dataset Versions**:\r\n   - There are three versions of the AmbigNQ dataset:\r\n     - **Light Version**: Contains only inputs and outputs.\r\n     - **Full Version**: Includes all annotation metadata.\r\n     - **Evidence Version**: Provides semi-oracle evidence articles along with questions and answers.\r\n\r\n(1) AmbigQA - University of Washington. https://nlp.cs.washington.edu/ambigqa/.\r\n(2) ambig_qa.py · ambig_qa at main - Hugging Face. https://huggingface.co/datasets/ambig_qa/blob/main/ambig_qa.py.\r\n(3) dataset_infos.json · ambig_qa at main - Hugging Face. https://huggingface.co/datasets/ambig_qa/blob/main/dataset_infos.json.\r\n(4) AmbigQA/AmbigNQ README - GitHub: Let’s build from here. https://github.com/shmsw25/AmbigQA.","description_withheld":null,"homepage":"https://nlp.cs.washington.edu/ambigqa","introduced_date":"2020-04-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/ambigqa-answering-ambiguous-open-domain","title":"AmbigQA: Answering Ambiguous Open-domain Questions","first_author":"Sewon Min","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["AmbigNQ"],"data_loaders":[],"num_papers_in_archive":11,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}