{"url":"/dataset/chinesesquad","name":"ChineseSquad","full_name":null,"description_markdown":"**ChineseSquad** (中文机器阅读理解数据集) is a dataset specifically designed for **Chinese machine reading comprehension**. It is created by translating and manually correcting the original **SQuAD (Stanford Question Answering Dataset)** into Chinese. The dataset includes both **V1.1** and **V2.0** versions of SQuAD. However, due to some translation challenges (especially with short answers and document translations), the Chinese version has slightly fewer examples compared to the original English SQuAD¹.\r\n\r\nHere are some key details about the ChineseSquad dataset:\r\n\r\n- **Data Sources**:\r\n    - ChineseSquad is derived from the original SQuAD dataset through machine translation and subsequent manual corrections.\r\n    - It includes both the **V1.1** and **V2.0** versions.\r\n- **Data Size**:\r\n    - The dataset contains both questions with answers and questions without answers.\r\n    - Here's a breakdown of the data:\r\n        - **squad-zen 1.0 train**: 68,213 examples with answers, 43,498 examples without answers (total: 110k)\r\n        - **squad-zen 1.0 dev**: 8,326 examples with answers, 5,954 examples without answers (total: 14k)\r\n        - **squad 2.0 train**: 46,530 examples with answers, 43,498 examples without answers (total: 90k)\r\n        - **squad 2.0 dev**: 3,391 examples with answers, 5,945 examples without answers (total: 9k)\r\n        - **squad 1.1 dev**: 7,679 examples (no answers provided)\r\n        - **squad 1.1 train**: 55,526 examples (no answers provided)","description_withheld":null,"homepage":"https://github.com/pluto-junzeng/ChineseSquad","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["ChineseSquad"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}