Datasets › WikiText-2

WikiText-2

Introduced by Stephen Merity et al. in Pointer Sentinel Mixture Models26 Sep 2016 archive 2025-07-28

The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.

Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far larger vocabulary and retains the original case, punctuation and numbers - all of which are removed in PTB. As it is composed of full articles, the dataset is well suited for models that can take advantage of long term dependencies.

Source: The WikiText Long Term Dependency Language Modeling Dataset Image Source: https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Language Modelling WikiText-2 SparseGPT (175B, 50% Sparsity) Test perplexity 8.21 SparseGPT: Massive Language Models Can Be Accurately... nvidia/tensorrt-model-optimizer +5 38 Compare

Papers archive 2025-07-28

23 shown of 23 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,081. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Advancing State of the Art in Language Modeling 1 1 28 Nov 2023 not harvested
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot 6 5 2 Jan 2023 ran 2 of 12 samples (10 unverified; 9 pointer-only for licence)
Efficient recurrent architectures through activity sparsity and sparse back-propagation through time 1 1 13 Jun 2022 not harvested
Hydra: A System for Large Multi-Model Deep Learning 1 1 16 Oct 2021 not harvested
Learning Associative Inference Using Fast Weight Memory 1 1 16 Nov 2020 ran 2 of 3 samples (1 unverified)
Alleviating Sequence Information Loss with Data Overlapping and Prime Batch Sizes 1 1 18 Sep 2019 not harvested
Mogrifier LSTM 3 2 4 Sep 2019 not harvested
Improving Neural Language Modeling via Adversarial Training 1 1 10 Jun 2019 not harvested
Deep Residual Output Layers for Neural Language Generation 1 2 14 May 2019 not harvested
Language Models with Transformers 1 1 20 Apr 2019 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Partially Shuffling the Training Data to Improve Language Models 1 2 11 Mar 2019 not harvested
Language Models are Unsupervised Multitask Learners 21 4 14 Feb 2019 not harvested
FRAGE: Frequency-Agnostic Word Representation 2 1 18 Sep 2018 not harvested
Direct Output Connection for a High-Rank Language Model 1 2 30 Aug 2018 not harvested
Improved Language Modeling by Decoding the Past 0 1 14 Aug 2018 not harvested
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model 9 2 10 Nov 2017 ran 1 of 23 samples (22 unverified)
Fraternal Dropout 1 1 31 Oct 2017 not harvested
Dynamic Evaluation of Neural Sequence Models 3 1 21 Sep 2017 ran 0 of 1 samples (1 unverified)
Gradual Learning of Recurrent Neural Networks 1 1 29 Aug 2017 not harvested
Regularizing and Optimizing LSTM Language Models 45 2 7 Aug 2017 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
On the State of the Art of Evaluation in Neural Language Models 1 1 18 Jul 2017 not harvested
Improving Neural Language Models with a Continuous Cache 14 2 13 Dec 2016 not harvested
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling 5 2 4 Nov 2016 not harvested

Dataset loaders archive 2025-07-28

9 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 3.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • WikiText-103
  • WikiText-2
  • wikitext wikitext-2-raw-v1

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections