{"url":"/dataset/faquad","name":"FaQuAD","full_name":null,"description_markdown":"The **FaQuAD** dataset is a **reading comprehension dataset** designed for evaluating question-answering models. Let me provide you with details about different versions of FaQuAD:\r\n\r\n1. **FaQuAD (English)**:\r\n    - **Description**: The original FaQuAD dataset follows the format of the **Stanford Question Answering Dataset (SQuAD)**.\r\n    - **Content**: It comprises **900 questions** related to **249 reading passages**. These passages were extracted from **18 official documents** of a computer science college at a Brazilian federal university and **21 Wikipedia articles** related to the Brazilian higher education system³.\r\n    - **Purpose**: Researchers use FaQuAD to develop and evaluate reading comprehension models in the domain of Brazilian higher education.\r\n    - **GitHub Repository**: You can find the dataset and related code on the [FaQuAD GitHub repository](https://github.com/liafacom/faquad).\r\n\r\n2. **FQuAD (French)**:\r\n    - **Description**: FQuAD is a **French Native Reading Comprehension dataset** created by higher education students. It consists of **25,000+ questions** based on a set of Wikipedia articles.\r\n    - **Similarity to SQuAD**: Like SQuAD, FQuAD provides annotated questions and answers for evaluation purposes.\r\n    - **Website**: You can explore the FQuAD dataset on the [FQuAD website](https://fquad.illuin.tech/).\r\n\r\n3. **FaQuAD (Portuguese)**:\r\n    - **Description**: As far as we know, FaQuAD is a **pioneer Portuguese reading comprehension dataset** that follows the challenging format of SQuAD.\r\n    - **Source**: The dataset includes passages from **official documents of a Brazilian computer science college** and **Wikipedia articles** related to Brazilian higher education.\r\n    - **GitHub Repository**: The FaQuAD data and source code for experiments are available on the [FaQuAD GitHub repository](https://github.com/liafacom/faquad).\r\n\r\nIn summary, FaQuAD provides valuable resources for training and evaluating question-answering models across different languages and domains. Researchers can use these datasets to advance natural language understanding and improve machine comprehension systems.\r\n\r\nSource: Conversation with Bing, 3/16/2024\r\n(1) FaQuAD: Reading Comprehension Dataset in the Domain of ... - ResearchGate. https://www.researchgate.net/profile/Eraldo-Fernandes/publication/337789791_FaQuAD_Reading_Comprehension_Dataset_in_the_Domain_of_Brazilian_Higher_Education/links/5e825f5fa6fdcc139c173c8f/FaQuAD-Reading-Comprehension-Dataset-in-the-Domain-of-Brazilian-Higher-Education.pdf.\r\n(2) GitHub - liafacom/faquad: FaQuAD reading comprehension dataset and .... https://github.com/liafacom/faquad.\r\n(3) FQuAD. https://fquad.illuin.tech/.\r\n(4) ruanchaves/faquad-nli · Datasets at Hugging Face. https://huggingface.co/datasets/ruanchaves/faquad-nli.\r\n(5) FaQuAD: Reading Comprehension Dataset in the Domain of ... - GitHub. https://github.com/liafacom/faquad?search=1.","description_withheld":null,"homepage":"https://github.com/liafacom/faquad/tree/master","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["FaQuAD"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}