Datasets › TyDiQA-GoldP
TyDiQA-GoldP
TyDiQA is the gold passage version of the Typologically Diverse Question Answering (TyDiWA) dataset, a benchmark for information-seeking question answering, which covers nine languages. The gold passage version is a simplified version of the primary task, which uses only the gold passage as context and excludes unanswerable questions. It is thus similar to XQuAD and MLQA, while being more challenging as questions have been written without seeing the answers, leading to 3× and 2× less lexical overlap compared to XQuAD and MLQA respectively.
Source: XTREME
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Cross-Lingual Question Answering | TyDiQA-GoldP | ByT5 (fine-tuned) EM 81.9 | ByT5: Towards a token-free future with pre-trained... | huggingface/transformers +4 | 11 | Compare |
Papers archive 2025-07-28
6 shown of 6 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 38. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| PaLM 2 Technical Report | 1 | 3 | 17 May 2023 | not harvested |
| Transcending Scaling Laws with 0.1% Extra Compute | 0 | 2 | 20 Oct 2022 | not harvested |
| Scaling Instruction-Finetuned Language Models | 9 | 2 | 20 Oct 2022 | ran 8 of 17 samples (9 unverified; 2 pointer-only for licence) |
| PaLM: Scaling Language Modeling with Pathways | 7 | 1 | 5 Apr 2022 | ran 30 of 37 samples (7 unverified) |
| ByT5: Towards a token-free future with pre-trained byte-to-byte models | 5 | 2 | 28 May 2021 | ran 0 of 6 samples (6 unverified) |
| Rethinking embedding coupling in pre-trained language models | 4 | 1 | 24 Oct 2020 | not harvested |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- TyDiQA-GoldP
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections