Datasets › Billion Word Benchmark

Billion Word Benchmark

Introduced by Ciprian Chelba et al. in One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling archive 2025-07-28

The One Billion Word dataset is a dataset for language modeling. The training/held-out data was produced from the WMT 2011 News Crawl data using a combination of Bash shell and Perl scripts.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Language Modelling One Billion Word MDLM (AR baseline) PPL 20.09 Simple and Effective Masked Diffusion Language Models kuleshov-group/mdlm +1 27 Compare
Text Generation One Billion Word WGANGP + DGflow JS-4 0.186 Refining Deep Generative Models via Discriminator Gradient Flow clear-nus/DGflow 1 Compare

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 141. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Simple and Effective Masked Diffusion Language Models 2 2 11 Jun 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences 2 2 25 Jul 2021 ran 5 of 10 samples (5 unverified)
OmniNet: Omnidirectional Representations from Transformers 1 3 1 Mar 2021 not harvested
When Attention Meets Fast Recurrence: Training Language Models with Reduced Compute 1 2 24 Feb 2021 not harvested
Refining Deep Generative Models via Discriminator Gradient Flow 1 1 1 Dec 2020 ran 0 of 1 samples (1 unverified)
Language Models are Unsupervised Multitask Learners 21 1 14 Feb 2019 not harvested
The Evolved Transformer 3 1 30 Jan 2019 not harvested
Pay Less Attention with Lightweight and Dynamic Convolutions 3 1 29 Jan 2019 not harvested
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context 37 2 9 Jan 2019 ran 63 of 143 samples (80 unverified; 43 pointer-only for licence)
Mesh-TensorFlow: Deep Learning for Supercomputers 1 1 5 Nov 2018 not harvested
Adaptive Input Representations for Neural Language Modeling 3 2 28 Sep 2018 not harvested
Factorization tricks for LSTM networks 2 1 31 Mar 2017 not harvested
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer 4 2 23 Jan 2017 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Language Modeling with Gated Convolutional Networks 11 1 23 Dec 2016 not harvested
Exploring the Limits of Language Modeling 10 3 7 Feb 2016 ran 2 of 8 samples (6 unverified)
Skip-gram Language Modeling Using Sparse Non-negative Matrix Probability Estimation 0 1 3 Dec 2014 not harvested
One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling 3 1 11 Dec 2013 not harvested

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Apache-2

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • One Billion Word
  • Billion Word Benchmark

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections