Datasets › LAMBADA

LAMBADA

Introduced by Denis Paperno et al. in The LAMBADA dataset: Word prediction requiring a broad discourse context1 Jan 2016 archive 2025-07-28

The LAMBADA (LAnguage Modeling Broadened to Account for Discourse Aspects) benchmark is an open-ended cloze task which consists of about 10,000 passages from BooksCorpus where a missing target word is predicted in the last sentence of each passage. The missing word is constrained to always be the last word of the last sentence and there are no candidate words to choose from. Examples were filtered by humans to ensure they were possible to guess given the context, i.e., the sentences in the passage leading up to the last sentence. Examples were further filtered to ensure that missing words could not be guessed without the context, ensuring that models attempting the dataset would need to reason over the entire paragraph to answer questions.

Source: Recent Advances in Natural Language Inference:A Survey of Benchmarks, Resources, and Approaches Image Source: https://arxiv.org/pdf/1606.06031.pdf

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Language Modelling LAMBADA PaLM-540B (Few-Shot) Accuracy 89.7 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 37 Compare

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 293. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Mamba: Linear-Time Sequence Modeling with Selective State Spaces 35 1 1 Dec 2023 ran 18 of 62 samples (44 unverified; 28 pointer-only for licence)
Stay on topic with Classifier-Free Guidance 0 3 30 Jun 2023 not harvested
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling 4 4 3 Apr 2023 not harvested
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot 6 5 2 Jan 2023 ran 2 of 12 samples (10 unverified; 9 pointer-only for licence)
GLM-130B: An Open Bilingual Pre-trained Model 9 1 5 Oct 2022 ran 5 of 21 samples (16 unverified)
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model 2 1 28 Jan 2022 not harvested
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts 0 1 13 Dec 2021 not harvested
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 8 2 18 Mar 2021 ran 1 of 1 samples (0 unverified)
Language Models are Few-Shot Learners 67 5 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Residual Shuffle-Exchange Networks for Fast Processing of Long Sequences 2 1 6 Apr 2020 not harvested
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts 3 1 29 Sep 2019 not harvested
Language Models are Unsupervised Multitask Learners 21 1 14 Feb 2019 not harvested
Universal Transformers 8 1 10 Jul 2018 ran 14 of 25 samples (11 unverified; 24 pointer-only for licence)
Broad Context Language Modeling as Reading Comprehension 0 1 26 Oct 2016 not harvested

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • LAMBADA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections