{"url":"/dataset/paq","name":"PAQ","full_name":"Probably Asked Questions","description_markdown":"**Probably Asked Questions** (**PAQ**) is a very large resource of 65M automatically-generated QA-pairs. PAQ is a semi-structured Knowledge Base (KB) of 65M natural language QA-pairs, which models can memorise and/or learn to retrieve from. PAQ differs from traditional KBs in that questions and answers are stored in natural language, and that questions are generated such that they are likely to appear in ODQA datasets. PAQ is automatically constructed using a question generation model and Wikipedia.\r\n\r\nSource: [Lewis et al.](https://arxiv.org/pdf/2102.07033.pdf)\r\n\r\nImage source: [Lewis et al.](https://arxiv.org/pdf/2102.07033.pdf)","description_withheld":null,"homepage":"https://github.com/facebookresearch/PAQ","introduced_date":"2021-02-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/paq-65-million-probably-asked-questions-and","title":"PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them","first_author":"Patrick Lewis","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["PAQ"],"data_loaders":[{"repo":"https://github.com/facebookresearch/PAQ","url":"https://github.com/facebookresearch/PAQ","frameworks":[]}],"num_papers_in_archive":53,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}