Datasets › Penn Treebank

Penn Treebank

Introduced in Building a Large Annotated Corpus of English: The Penn Treebank1 Jan 1993 archive 2025-07-28

The English Penn Treebank (PTB) corpus, and in particular the section of the corpus corresponding to the articles of Wall Street Journal (WSJ), is one of the most known and used corpus for the evaluation of models for sequence labelling. The task consists of annotating each word with its Part-of-Speech tag. In the most common split of this corpus, sections from 0 to 18 are used for training (38 219 sentences, 912 344 tokens), sections from 19 to 21 are used for validation (5 527 sentences, 131 768 tokens), and sections from 22 to 24 are used for testing (5 462 sentences, 129 654 tokens). The corpus is also commonly used for character-level and word-level Language Modelling.

Source: Seq2Biseq: Bidirectional Output-wise Recurrent Neural Networks for Sequence Modelling Image Source: https://dl.acm.org/doi/10.5555/972470.972475

Benchmarks archive 2025-07-28

All 10 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Language Modelling Penn Treebank (Word Level) GPT-3 (Zero-Shot) Test perplexity 20.5 Language Models are Few-Shot Learners ggml-org/llama.cpp +66 43 Compare
Constituency Parsing Penn Treebank Hashing + XLNet F1 score 96.43 To be Continuous, or to be Discrete, Those are Bits of Questions speedcell4/parserker 27 Compare
Dependency Parsing Penn Treebank Label Attention Layer + HPSG + XLNet LAS 96.26 Rethinking Self-Attention: Towards Interpretability in... KhalilMrini/LAL-Parser +1 22 Compare
Language Modelling Penn Treebank (Character Level) Mogrifier LSTM + dynamic eval Bit per Character (BPC) 1.083 Mogrifier LSTM deepmind/lamb +2 20 Compare
Part-Of-Speech Tagging Penn Treebank SALE-BART encoder Accuracy 98.15 Sequence Alignment Ensemble with a Single Neural Network... — 20 Compare
Chunking Penn Treebank ACE F1 score 97.3 Automated Concatenation of Embeddings for Structured Prediction Alibaba-NLP/ACE +1 8 Compare
Unsupervised Dependency Parsing Penn Treebank Ensemble (selected w/ society entropy) UAS 67.3 Error Diversity Matters: An Error-Resistant Ensemble... manga-uofa/ed4udp 6 Compare
Open Information Extraction Penn Treebank Deepstruct zero-shot F1 51 DeepStruct: Pretraining of Language Models for Structure... cgraywang/deepstruct 4 Compare
Stochastic Optimization Penn Treebank (Character Level) 3x1000 LSTM - 500 Epochs AvaGrad Bit per Character (BPC) 1.175 Domain-independent Dominance of Adaptive Methods lolemacs/avagrad 4 Compare
Constituency Grammar Induction Penn Treebank no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 108 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,006. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing 1 1 16 Dec 2024 not harvested
To be Continuous, or to be Discrete, Those are Bits of Questions 1 2 12 Jun 2024 not harvested
Advancing State of the Art in Language Modeling 1 1 28 Nov 2023 not harvested
Enhancing Structure-aware Encoder with Extremely Limited Data for Graph-based Dependency Parsing 1 1 1 Oct 2022 not harvested
Sequence Alignment Ensemble with a Single Neural Network for Sequence Labeling 0 1 7 Jul 2022 not harvested
DeepStruct: Pretraining of Language Models for Structure Prediction 1 3 21 May 2022 ran 7 of 13 samples (6 unverified)
Sequential Alignment Methods for Ensemble Part-of-Speech Tagging 1 1 23 Mar 2022 not harvested
Investigating Non-local Features for Neural Constituency Parsing 1 1 27 Sep 2021 not harvested
Zero-Shot Information Extraction as a Unified Text-to-Triple Translation 1 1 23 Sep 2021 not harvested
N-ary Constituent Tree Parsing with Recursive Semi-Markov Model 1 1 26 Jul 2021 not harvested
AutoDropout: Learning Dropout Patterns to Regularize Deep Networks 1 1 5 Jan 2021 not harvested
Learning Associative Inference Using Fast Weight Memory 1 1 16 Nov 2020 ran 2 of 3 samples (1 unverified)
Strongly Incremental Constituency Parsing with Graph Neural Networks 3 1 27 Oct 2020 ran 2 of 5 samples (3 unverified)
Improving Constituency Parsing with Span Attention 1 1 15 Oct 2020 ran 1 of 1 samples (0 unverified)
Second-Order Neural Dependency Parsing with Message Passing and End-to-End Training 1 1 10 Oct 2020 not harvested
Automated Concatenation of Embeddings for Structured Prediction 2 2 10 Oct 2020 not harvested
Fast and Accurate Neural CRF Constituency Parsing 2 3 9 Aug 2020 ran 3 of 3 samples (0 unverified)
Language Models are Few-Shot Learners 67 1 28 May 2020 ran 15 of 65 samples (50 unverified; 4 pointer-only for licence)
Efficient Second-Order TreeCRF for Neural Dependency Parsing 2 1 3 May 2020 ran 2 of 3 samples (1 unverified)
Recursive Non-Autoregressive Graph-to-Graph Transformer for Dependency Parsing with Iterative Refinement 1 1 29 Mar 2020 not harvested
Addressing Some Limitations of Transformers with Feedback Memory 4 1 21 Feb 2020 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)
Recurrent Highway Networks with Grouped Auxiliary Memory 4 1 13 Dec 2019 not harvested
Domain-independent Dominance of Adaptive Methods 1 4 4 Dec 2019 not harvested
Gating Revisited: Deep Multi-layer RNNs That Can Be Trained 3 1 25 Nov 2019 not harvested
Seq-U-Net: A One-Dimensional Causal U-Net for Efficient Sequence Modelling 1 4 14 Nov 2019 ran 0 of 6 samples (6 unverified)
Rethinking Self-Attention: Towards Interpretability in Neural Parsing 2 2 10 Nov 2019 ran 1 of 3 samples (2 unverified; 3 pointer-only for licence)
Generalizing Natural Language Analysis through Span-relation Representations 3 3 10 Nov 2019 not harvested
Deep Independently Recurrent Neural Network (IndRNN) 1 3 11 Oct 2019 not harvested
Mogrifier LSTM 3 3 4 Sep 2019 not harvested
Deep Equilibrium Models 11 1 3 Sep 2019 ran 3 of 12 samples (9 unverified; 4 pointer-only for licence)

The full list of 108 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Penn Treebank
  • Penn Treebank (Word Level)
  • Penn Treebank (Character Level)
  • Penn Treebank (Character Level) 3x1000 LSTM - 500 Epochs

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections