Datasets › WikiText-103

WikiText-103

Introduced by Stephen Merity et al. in Pointer Sentinel Mixture Models26 Sep 2016 archive 2025-07-28

The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License.

Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far larger vocabulary and retains the original case, punctuation and numbers - all of which are removed in PTB. As it is composed of full articles, the dataset is well suited for models that can take advantage of long term dependencies.

Source: The WikiText Long Term Dependency Language Modeling Dataset Image Source: https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Language Modelling WikiText-103 RETRO (7.5B) Test perplexity 2.4 Improving language models by retrieving from trillions of tokens labmlai/annotated_deep_learning_paper_implementations +1 89 Compare
Text Generation WikiText-103 no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 55 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 560. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Advancing State of the Art in Language Modeling 1 1 28 Nov 2023 not harvested
Memory-efficient Stochastic methods for Memory-based Transformers 1 1 14 Nov 2023 not harvested
GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling 3 1 3 Nov 2023 ran 3 of 5 samples (2 unverified)
The Information Pathways Hypothesis: Transformers are Dynamic Self-Ensembles 1 2 2 Jun 2023 not harvested
Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal Representation 1 1 31 May 2023 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Hyena Hierarchy: Towards Larger Convolutional Language Models 7 2 21 Feb 2023 ran 5 of 5 samples (0 unverified; 4 pointer-only for licence)
Hungry Hungry Hippos: Towards Language Modeling with State Space Models 3 5 28 Dec 2022 ran 7 of 15 samples (8 unverified)
You can't pick your neighbors, or can you? When and how to rely on retrieval in the $k$NN-LM 1 1 28 Oct 2022 ran 0 of 3 samples (3 unverified; 3 pointer-only for licence)
Mega: Moving Average Equipped Gated Attention 7 1 21 Sep 2022 ran 12 of 13 samples (1 unverified; 8 pointer-only for licence)
General-purpose, long-context autoregressive modeling with Perceiver AR 3 1 15 Feb 2022 ran 6 of 12 samples (6 unverified; 1 pointer-only for licence)
Improving language models by retrieving from trillions of tokens 2 1 8 Dec 2021 ran 16 of 23 samples (7 unverified; 3 pointer-only for licence)
Efficiently Modeling Long Sequences with Structured State Spaces 8 1 31 Oct 2021 ran 28 of 55 samples (27 unverified; 3 pointer-only for licence)
∞-former: Infinite Memory Transformer 1 7 1 Sep 2021 not harvested
FNetAR: Mixing Tokens with Autoregressive Fourier Transforms 1 1 22 Jul 2021 not harvested
Differentiable Model Compression via Pseudo Quantization Noise 1 1 20 Apr 2021 not harvested
Revisiting Simple Neural Probabilistic Language Models 1 1 8 Apr 2021 ran 0 of 3 samples (3 unverified)
Finetuning Pretrained Transformers into RNNs 2 1 24 Mar 2021 ran 3 of 8 samples (5 unverified; 4 pointer-only for licence)
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 8 2 18 Mar 2021 ran 1 of 1 samples (0 unverified)
Random Feature Attention 0 2 3 Mar 2021 not harvested
When Attention Meets Fast Recurrence: Training Language Models with Reduced Compute 1 2 24 Feb 2021 not harvested
Subformer: A Parameter Reduced Transformer 0 1 1 Jan 2021 not harvested
Shortformer: Better Language Modeling using Shorter Inputs 1 2 31 Dec 2020 not harvested
Rethinking Attention with Performers 7 1 30 Sep 2020 ran 9 of 16 samples (7 unverified; 6 pointer-only for licence)
Pay Attention when Required 2 2 9 Sep 2020 not harvested
DeLighT: Deep and Light-weight Transformer 2 1 3 Aug 2020 ran 0 of 3 samples (3 unverified)
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention 8 1 29 Jun 2020 ran 3 of 8 samples (5 unverified; 3 pointer-only for licence)
How much complexity does an RNN architecture need to learn syntax-sensitive dependencies? 1 3 17 May 2020 not harvested
Segatron: Segment-Aware Transformer for Language Modeling and Understanding 1 1 30 Apr 2020 not harvested
Efficient Content-Based Sparse Attention with Routing Transformers 2 1 12 Mar 2020 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)
Addressing Some Limitations of Transformers with Feedback Memory 4 2 21 Feb 2020 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)

The full list of 55 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-SA 3.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • WikiText-103

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections