Datasets › TriviaQA

TriviaQA

Introduced by Mandar Joshi et al. in TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension1 Jan 2017 archive 2025-07-28

TriviaQA is a realistic text-based question answering dataset which includes 950K question-answer pairs from 662K documents collected from Wikipedia and the web. This dataset is more challenging than standard QA benchmark datasets such as Stanford Question Answering Dataset (SQuAD), as the answers for a question may not be directly obtained by span prediction and the context is very long. TriviaQA dataset consists of both human-verified and machine-generated QA subsets.

Source: Episodic Memory Reader: Learning What to Rememberfor Question Answering from Streaming Data Image Source: Joshi et al

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering TriviaQA Claude 2 (few-shot, k=5) EM 87.5 Model Card and Evaluations for Claude Models — 56 Compare
Open-Domain Question Answering KILT: TriviaQA Re2G KILT-EM 57.91 Re2G: Retrieve, Rerank, Generate ibm/kgi-slot-filling 15 Compare
Question Generation TriviaQA Info-HCVAE QAE 35.45 Generating Diverse and Consistent QA pairs from Contexts... seanie12/Info-HCVAE 2 Compare
Open-Domain Question Answering TriviaQA UnitedQA (Hybrid) Exact Match 70.5 UnitedQA: A Hybrid Approach for Open Domain Question Answering — 1 Compare
Text Generation TriviaQA no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 41 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 953. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Search-o1: Agentic Search-Enhanced Large Reasoning Models 2 1 9 Jan 2025 not harvested
SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments 0 1 15 Oct 2024 not harvested
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs 0 3 2 Jul 2024 not harvested
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation 1 1 26 Jun 2024 ran 1 of 3 samples (2 unverified; 3 pointer-only for licence)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling 1 1 18 Jun 2024 ran 4 of 7 samples (3 unverified)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 1 1 12 Mar 2024 not harvested
ChatQA: Surpassing GPT-4 on Conversational QA and RAG 0 3 18 Jan 2024 not harvested
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
RA-DIT: Retrieval-Augmented Dual Instruction Tuning 0 1 2 Oct 2023 not harvested
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 1 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
Model Card and Evaluations for Claude Models 0 3 11 Jul 2023 not harvested
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
GPT-4 Technical Report 11 1 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
REPLUG: Retrieval-Augmented Black-Box Language Models 3 2 30 Jan 2023 ran 0 of 13 samples (13 unverified; 13 pointer-only for licence)
FiE: Building a Global Probability Space by Leveraging Early Fusion in Encoder for Open-Domain Question Answering 0 1 18 Nov 2022 not harvested
DyREx: Dynamic Query Representation for Extractive Question Answering 1 1 26 Oct 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
Re2G: Retrieve, Rerank, Generate 1 1 13 Jul 2022 ran 1 of 8 samples (7 unverified)
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 ran 30 of 37 samples (7 unverified)
LinkBERT: Pretraining Language Models with Document Links 1 1 29 Mar 2022 ran 0 of 14 samples (14 unverified)
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts 0 3 13 Dec 2021 not harvested
Mention Memory: incorporating textual knowledge into Transformers through entity mention attention 1 1 12 Oct 2021 not harvested
ReasonBERT: Pre-trained to Reason with Distant Supervision 1 2 10 Sep 2021 ran 3 of 11 samples (8 unverified)
Finetuned Language Models Are Zero-Shot Learners 8 1 3 Sep 2021 ran 0 of 1 samples (1 unverified)
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering 2 1 9 Jun 2021 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
UnitedQA: A Hybrid Approach for Open Domain Question Answering 0 2 1 Jan 2021 not harvested
Distilling Knowledge from Reader to Retriever for Question Answering 4 1 8 Dec 2020 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
KILT: a Benchmark for Knowledge Intensive Language Tasks 3 1 4 Sep 2020 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering 8 1 2 Jul 2020 not harvested

The full list of 41 is in the JSON twin.

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • TriviaQA
  • KILT: TriviaQA

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections