Datasets › KILT

KILT (KILT Benchmark)

Introduced by Fabio Petroni et al. in KILT: a Benchmark for Knowledge Intensive Language Tasks4 Sep 2020 archive 2025-07-28

KILT (Knowledge Intensive Language Tasks) is a benchmark consisting of 11 datasets representing 5 types of tasks:

  • Fact-checking (FEVER),
  • Entity linking (AIDA CoNLL-YAGO, WNED-WIKI, WNED-CWEB),
  • Slot filling (T-Rex, Zero Shot RE),
  • Open domain QA (Natural Questions, HotpotQA, TriviaQA, ELI5),
  • Dialog generation (Wizard of Wikipedia).

All these datasets have been grounded in a single pre-processed wikipedia snapshot, allowing for fairer and more consistent evaluation as well as enabling new task setups such as multitask and transfer learning.

Source: KILT Benchmarking

Benchmarks archive 2025-07-28

All 11 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Fact Verification KILT: FEVER Re2G KILT-AC 78.53 Re2G: Retrieve, Rerank, Generate ibm/kgi-slot-filling 33 Compare
Open-Domain Dialog KILT: Wizard of Wikipedia Hindsight KILT-RL 11.92 — — 21 Compare
Slot Filling KILT: Zero Shot RE single ngram KILT-AC 73.2 — — 21 Compare
Slot Filling KILT: T-REx Re2G KILT-AC 75.84 Re2G: Retrieve, Rerank, Generate ibm/kgi-slot-filling 20 Compare
Open-Domain Question Answering KILT: Natural Questions Re2G KILT-EM 43.56 Re2G: Retrieve, Rerank, Generate ibm/kgi-slot-filling 16 Compare
Open-Domain Question Answering KILT: ELI5 somebody KILT-RL 2.62 — — 16 Compare
Open-Domain Question Answering KILT: HotpotQA intersect KILT-EM 18.06 — — 14 Compare
Entity Linking KILT: AIDA-YAGO2 GENRE KILT-AC 89.85 Autoregressive Entity Retrieval facebookresearch/GENRE +1 11 Compare
Entity Linking KILT: WNED-CWEB GENRE KILT-AC 71.22 Autoregressive Entity Retrieval facebookresearch/GENRE +1 10 Compare
Entity Linking KILT: WNED-WIKI GENRE KILT-AC 87.44 Autoregressive Entity Retrieval facebookresearch/GENRE +1 10 Compare
Question Answering KILT: ELI5 RBG Rouge-L 27.13 Read before Generate! Faithful Long Form Question... — 7 Compare

Papers archive 2025-07-28

8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 117. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
An Efficient Memory-Augmented Transformer for Knowledge-Intensive NLP Tasks 1 1 30 Oct 2022 not harvested
Re2G: Retrieve, Rerank, Generate 1 4 13 Jul 2022 ran 1 of 8 samples (7 unverified)
Knowledge Infused Decoding 1 1 6 Apr 2022 ran 0 of 3 samples (3 unverified)
Read before Generate! Faithful Long Form Question Answering with Machine Reading 0 1 1 Mar 2022 not harvested
Hurdles to Progress in Long-form Question Answering 2 2 10 Mar 2021 not harvested
Learning Dense Representations of Phrases at Scale 4 2 23 Dec 2020 not harvested
Autoregressive Entity Retrieval 2 3 2 Oct 2020 not harvested
KILT: a Benchmark for Knowledge Intensive Language Tasks 3 14 4 Sep 2020 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • NQ KILT
  • KILT: Zero Shot RE
  • KILT: Wizard of Wikipedia
  • KILT: WNED-WIKI
  • KILT: WNED-CWEB
  • KILT: TriviaQA
  • KILT: T-REx
  • KILT: Natural Questions
  • KILT: HotpotQA
  • KILT: FEVER
  • KILT: ELI5
  • KILT: AIDA-YAGO2
  • KILT

13 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections