Datasets › Text8

Text8

archive 2025-07-28

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Language Modelling Text8 GPT-2 Bit per Character (BPC) 0.98 Language Models are Unsupervised Multitask Learners huggingface/transformers +20 24 Compare

Papers archive 2025-07-28

19 shown of 19 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 22. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Bayesian Flow Networks 1 1 14 Aug 2023 ran 11 of 14 samples (3 unverified)
Focus Your Attention (with Adaptive IIR Filters) 0 1 24 May 2023 not harvested
Long-Short Transformer: Efficient Transformers for Language and Vision 3 1 5 Jul 2021 ran 4 of 4 samples (0 unverified; 2 pointer-only for licence)
Pay Attention when Required 2 1 9 Sep 2020 not harvested
Recurrent Highway Networks with Grouped Auxiliary Memory 4 1 13 Dec 2019 not harvested
BP-Transformer: Modelling Long-Range Context via Binary Partitioning 2 1 11 Nov 2019 ran 0 of 7 samples (7 unverified)
Augmenting Self-attention with Persistent Memory 2 2 2 Jul 2019 ran 5 of 5 samples (0 unverified; 4 pointer-only for licence)
Discrete Flows: Invertible Generative Models of Discrete Data 2 1 24 May 2019 not harvested
Adaptive Attention Span in Transformers 8 2 19 May 2019 not harvested
Dynamic Evaluation of Transformer Language Models 1 1 17 Apr 2019 not harvested
Language Models are Unsupervised Multitask Learners 21 1 14 Feb 2019 not harvested
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context 37 1 9 Jan 2019 ran 63 of 143 samples (80 unverified; 43 pointer-only for licence)
Character-Level Language Modeling with Deeper Self-Attention 1 2 9 Aug 2018 not harvested
Dynamic Evaluation of Neural Sequence Models 3 1 21 Sep 2017 ran 0 of 1 samples (1 unverified)
Multiplicative LSTM for sequence modelling 1 2 26 Sep 2016 not harvested
Hierarchical Multiscale Recurrent Neural Networks 3 1 6 Sep 2016 not harvested
Recurrent Highway Networks 6 1 12 Jul 2016 ran 1 of 5 samples (4 unverified)
Recurrent Batch Normalization 3 1 30 Mar 2016 ran 0 of 2 samples (2 unverified)
Architectural Complexity Measures of Recurrent Neural Networks 0 2 26 Feb 2016 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Text8

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections