Datasets › IMDb Movie Reviews

IMDb Movie Reviews

Introduced in Learning Word Vectors for Sentiment Analysis1 Jan 2011 archive 2025-07-28

The IMDb Movie Reviews dataset is a binary sentiment analysis dataset consisting of 50,000 reviews from the Internet Movie Database (IMDb) labeled as positive or negative. The dataset contains an even number of positive and negative reviews. Only highly polarizing reviews are considered. A negative review has a score ≤ 4 out of 10, and a positive review has a score ≥ 7 out of 10. No more than 30 reviews are included per movie. The dataset contains additional unlabeled data.

Source: http://nlpprogress.com/english/sentiment_analysis.html Image Source: Maas et al

Benchmarks archive 2025-07-28

All 9 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 62 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,787. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Learning in Wilson-Cowan model for metapopulation 1 1 24 Jun 2024 not harvested
LlamBERT: Large-scale low-cost data annotation in NLP 1 3 23 Mar 2024 not harvested
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs 1 5 30 Jan 2024 not harvested
Cache me if you Can: an Online Cost-aware Teacher-Student framework to Reduce the Calls to Large Language Models 1 1 20 Oct 2023 ran 1 of 1 samples (0 unverified)
SplitEE: Early Exit in Deep Neural Networks with Split Computing 1 1 17 Sep 2023 not harvested
Analysis of the Evolution of Advanced Transformer-Based Language Models: Experiments on Opinion Mining 1 15 7 Aug 2023 not harvested
An Algorithm for Routing Vectors in Sequences 1 2 20 Nov 2022 not harvested
The Document Vectors Using Cosine Similarity Revisited 1 4 26 May 2022 not harvested
Context-Aware Compilation of DNN Training Pipelines across Edge and Cloud 1 1 30 Dec 2021 not harvested
Finetuned Language Models Are Zero-Shot Learners 8 2 3 Sep 2021 ran 0 of 1 samples (1 unverified)
MA-BERT: Learning Representation by Incorporating Multi-Attribute Knowledge in Transformers 1 1 1 Aug 2021 not harvested
Closed-form Continuous-time Neural Models 1 1 25 Jun 2021 ran 2 of 10 samples (8 unverified)
Classifying Textual Data with Pre-trained Vision Models through Transfer Learning and Data Transformations 1 3 23 Jun 2021 not harvested
Entailment as Few-Shot Learner 3 1 29 Apr 2021 ran 1 of 3 samples (2 unverified)
UnICORNN: A recurrent model for learning very long time dependencies 1 1 9 Mar 2021 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Improving Document-Level Sentiment Classification Using Importance of Sentences 0 1 9 Mar 2021 not harvested
Parallelizing Legendre Memory Unit Training 2 1 22 Feb 2021 not harvested
Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention 10 1 7 Feb 2021 ran 1 of 2 samples (1 unverified; 1 pointer-only for licence)
ERNIE-Doc: A Retrospective Long-Document Modeling Transformer 2 2 31 Dec 2020 not harvested
Neural Semi-supervised Learning for Text Classification Under Large-Scale Pretraining 1 1 17 Nov 2020 not harvested
Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training Approach 1 1 15 Oct 2020 ran 0 of 5 samples (5 unverified)
Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependencies 1 1 2 Oct 2020 ran 1 of 3 samples (2 unverified)
Revisiting LSTM Networks for Semi-Supervised Text Classification via Mixed Objective Function 1 1 8 Sep 2020 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
BP-Transformer: Modelling Long-Range Context via Binary Partitioning 2 1 11 Nov 2019 ran 0 of 7 samples (7 unverified)
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter 37 1 2 Oct 2019 ran 19 of 27 samples (8 unverified)
Rethinking Attribute Representation and Injection for Sentiment Classification 1 1 26 Aug 2019 not harvested
Message Passing Attention Networks for Document Understanding 2 1 17 Aug 2019 not harvested
Sentiment Classification Using Document Embeddings Trained with Cosine Similarity 2 1 1 Jul 2019 not harvested
Graph Star Net for Generalized Multi-Task Learning 1 1 21 Jun 2019 ran 1 of 6 samples (5 unverified)
XLNet: Generalized Autoregressive Pretraining for Language Understanding 27 1 19 Jun 2019 ran 10 of 24 samples (14 unverified; 3 pointer-only for licence)

The full list of 62 is in the JSON twin.

Dataset loaders archive 2025-07-28

21 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • IMDb
  • IMDb Movie Reviews
  • User and product information

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections