{"url":"/dataset/sentiment-merged","name":"Sentiment Merged","full_name":"SST-3, DynaSent R1/R2","description_markdown":"This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive). It is a merge of [Stanford Sentiment Treebank](https://nlp.stanford.edu/sentiment/) (SST-3) and [DynaSent](https://github.com/cgpotts/dynasent) Rounds 1 and 2, licensed under Apache 2.0 and Creative Commons Attribution 4.0 respectively. The SST-3, DynaSent R1, and DynaSent R2 datasets were randomly mixed to form a new dataset with 102,097 Train examples, 5,421 Validation examples, and 6,530 Test examples. See Table 1 for the distribution of labels within this merged dataset.\r\n\r\n### Table 1: Label Distribution for the Merged Dataset\r\n\r\n| Split      | Negative | Neutral | Positive |\r\n|------------|----------|---------|----------|\r\n| Train      | 21,910   | 49,148  | 31,039   |\r\n| Validation | 1,868    | 1,669   | 1,884    |\r\n| Test       | 2,352    | 1,829   | 2,349    |\r\n\r\n### Table 2: Contribution of Sources to the Merged Dataset\r\n\r\n| Dataset           | Samples | Percent (%) |\r\n|-------------------|---------|-------------|\r\n| DynaSent R1 Train | 80,488  | 78.83       |\r\n| DynaSent R2 Train | 13,065  | 12.80       |\r\n| SST-3 Train       | 8,544   | 8.37        |\r\n| **Total**         | 102,097 | 100.00      |\r\n\r\n### Source Datasets\r\n\r\nSST-5 is the [Stanford Sentiment Treebank](https://huggingface.co/datasets/stanfordnlp/sst) 5-way classification (positive, somewhat positive, neutral, somewhat negative, negative). To create SST-3 (positive, neutral, negative), the 'somewhat positive' class was merged and treated as 'positive'. Similarly, the 'somewhat negative class' was merged and treated as 'negative'.\r\n\r\n[DynaSent](https://huggingface.co/datasets/dynabench/dynasent) is a sentiment analysis dataset and dynamic benchmark with three classification labels: positive, negative, and neutral. The dataset was created in two rounds. First, a RoBERTa model was fine-tuned on a variety of datasets including SST-3, IMBD, and Yelp. They then extracted challenging sentences that fooled the model, and validated them with humans. For Round 2, a new RoBERTa model was trained on similar (but different) data, including the Round 1 dataset. The Dynabench platform was then used to create sentences written by workers that fooled the model.\r\n\r\nIt’s worth noting that the source datasets all have class imbalances. SST-3 positive and negative are about twice the number of neutral. In DynaSent R1, the neutral are more than three times the negative. And in DynaSent R2, the positive are more than double the neutral. Although this imbalance may be by design for DynaSent (to focus on the more challenging neutral class), it still represents an imbalanced dataset. The risk is that the model will learn mostly the dominant class.\r\nMerging the data helps mitigate this imbalance. Although there is still a majority of neutral examples in the training dataset, the neutral to negative ratio in DynaSent R1 is 3.21, and this is improved to 2.24 in the merged dataset.\r\n\r\nAnother potential issue is that the models will learn the dominant dataset, which is DynaSent R1. See Table 2 for a breakdown of the sources contributing to the Merged dataset.","description_withheld":null,"homepage":"https://huggingface.co/datasets/jbeno/sentiment_merged","introduced_date":"2024-12-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/electra-and-gpt-4o-cost-effective-partners","title":"ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis","first_author":"James P. Beno","url":null},"license":{"name":"MIT","url":"https://huggingface.co/datasets/jbeno/sentiment_merged/blob/main/README.md"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Sentiment Analysis","url":"/task/sentiment-analysis","datasets_with_task":"/datasets/task/sentiment-analysis"},{"name":"Sentiment Classification","url":"/task/sentiment-classification","datasets_with_task":"/datasets/task/sentiment-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Sentiment Merged"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/sentiment-analysis-on-sentiment-merged","task":"Sentiment Analysis","dataset_variant":"Sentiment Merged","rows":10,"metrics":["Macro F1"],"first_row_in_archive_order":{"model":"GPT-4o Fine-Tuned (Minimal)","paper":"/paper/electra-and-gpt-4o-cost-effective-partners","metrics":{"Macro F1":"86.99"},"code_links":[{"title":"jbeno/sentiment","url":"https://github.com/jbeno/sentiment"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/electra-and-gpt-4o-cost-effective-partners","title":"ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis","date":"2024-12-29","rows_on_this_dataset":10,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}