Datasets › SST

SST (Stanford Sentiment Treebank)

Introduced by Richard Socher et al. in Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank1 Oct 2013 archive 2025-07-28

The Stanford Sentiment Treebank is a corpus with fully labeled parse trees that allows for a complete analysis of the compositional effects of sentiment in language. The corpus is based on the dataset introduced by Pang and Lee (2005) and consists of 11,855 single sentences extracted from movie reviews. It was parsed with the Stanford parser and includes a total of 215,154 unique phrases from those parse trees, each annotated by 3 human judges.

Each phrase is labelled as either negative, somewhat negative, neutral, somewhat positive or positive. The corpus with all 5 labels is referred to as SST-5 or SST fine-grained. Binary classification experiments on full sentences (negative or somewhat negative vs somewhat positive or positive with neutral sentences discarded) refer to the dataset as SST-2 or SST binary.

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 84 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 2,354. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CAPO: Cost-Aware Prompt Optimization 2 3 22 Apr 2025 not harvested
SubRegWeigh: Effective and Efficient Annotation Weighing with Subword Regularization 1 1 10 Sep 2024 not harvested
SAME: Uncovering GNN Black Box with Structure-aware Shapley-based Multipiece Explanations 1 1 21 Sep 2023 not harvested
LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot Fine-Tuning 1 2 29 May 2023 not harvested
An Algorithm for Routing Vectors in Sequences 1 2 20 Nov 2022 not harvested
Fine-mixing: Mitigating Backdoors in Fine-tuned Language Models 1 1 18 Oct 2022 ran 0 of 19 samples (19 unverified; 19 pointer-only for licence)
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale 4 1 15 Aug 2022 ran 2 of 5 samples (3 unverified)
Adversarial Self-Attention for Language Understanding 1 2 25 Jun 2022 not harvested
Dual Contrastive Learning: Text Classification via Label-Aware Data Augmentation 2 1 21 Jan 2022 not harvested
Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners 4 1 30 Aug 2021 ran 3 of 4 samples (1 unverified)
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization 2 1 23 Jun 2021 ran 7 of 10 samples (3 unverified)
Pay Attention to MLPs 20 1 17 May 2021 ran 34 of 44 samples (10 unverified; 11 pointer-only for licence)
An Effective Baseline for Robustness to Distributional Shift 1 1 15 May 2021 not harvested
FNet: Mixing Tokens with Fourier Transforms 12 1 9 May 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
Entailment as Few-Shot Learner 3 1 29 Apr 2021 ran 1 of 3 samples (2 unverified)
How to Train BERT with an Academic Budget 4 1 15 Apr 2021 not harvested
Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention 10 1 7 Feb 2021 ran 1 of 2 samples (1 unverified; 1 pointer-only for licence)
Muppet: Massive Multi-task Representations with Pre-Finetuning 2 2 26 Jan 2021 not harvested
CLEAR: Contrastive Learning for Sentence Representation 0 1 31 Dec 2020 not harvested
RealFormer: Transformer Likes Residual Attention 5 1 21 Dec 2020 not harvested
Self-Explaining Structures Improve NLP Models 1 1 3 Dec 2020 not harvested
A Statistical Framework for Low-bitwidth Training of Deep Neural Networks 2 1 27 Oct 2020 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
Pay Attention when Required 2 1 9 Sep 2020 not harvested
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
SqueezeBERT: What can computer vision teach NLP about efficient neural networks? 6 1 19 Jun 2020 ran 0 of 1 samples (1 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators 19 1 23 Mar 2020 ran 26 of 40 samples (14 unverified; 10 pointer-only for licence)
Learning to Encode Position for Transformer with Continuous Dynamical Model 1 1 13 Mar 2020 ran 3 of 6 samples (3 unverified; 6 pointer-only for licence)
BP-Transformer: Modelling Long-Range Context via Binary Partitioning 2 1 11 Nov 2019 ran 0 of 7 samples (7 unverified)
SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization 6 6 8 Nov 2019 ran 6 of 8 samples (2 unverified; 1 pointer-only for licence)

The full list of 84 is in the JSON twin.

Dataset loaders archive 2025-07-28

7 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • SST-2
  • SST-2 Binary classification Dev
  • SST-5
  • sst2-es-mt
  • SST2
  • SST-2 Binary classification
  • SST-5 Fine-grained classification
  • SST

8 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections