Datasets › Hutter Prize

Hutter Prize

archive 2025-07-28

The Hutter Prize Wikipedia dataset, also known as enwiki8, is a byte-level dataset consisting of the first 100 million bytes of a Wikipedia XML dump. For simplicity we shall refer to it as a character-level dataset. Within these 100 million bytes are 205 unique tokens.

Source: NLP Progress

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Language Modelling Hutter Prize Transformer-XL + RMS dynamic eval Bit per Character (BPC) 0.94 Dynamic Evaluation of Transformer Language Models benkrause/dynamiceval-transformer 18 Compare

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 12. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Longformer: The Long-Document Transformer 22 2 10 Apr 2020 ran 15 of 35 samples (20 unverified; 5 pointer-only for licence)
Compressive Transformers for Long-Range Sequence Modelling 6 1 13 Nov 2019 ran 3 of 11 samples (8 unverified)
Mogrifier LSTM 3 2 4 Sep 2019 not harvested
Dynamic Evaluation of Transformer Language Models 1 1 17 Apr 2019 not harvested
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context 37 3 9 Jan 2019 ran 63 of 143 samples (80 unverified; 43 pointer-only for licence)
Character-Level Language Modeling with Deeper Self-Attention 1 2 9 Aug 2018 not harvested
An Analysis of Neural Language Modeling at Multiple Scales 12 1 22 Mar 2018 ran 3 of 18 samples (15 unverified)
Dynamic Evaluation of Neural Sequence Models 3 1 21 Sep 2017 ran 0 of 1 samples (1 unverified)
Fast-Slow Recurrent Neural Networks 1 2 24 May 2017 not harvested
Multiplicative LSTM for sequence modelling 1 1 26 Sep 2016 not harvested
Recurrent Highway Networks 6 2 12 Jul 2016 ran 1 of 5 samples (4 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Hutter Prize

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections