{"url":"/dataset/multilingual-sentiment-datasets","name":"Multilingual Sentiment Datasets","full_name":null,"description_markdown":"A collection of multilingual sentiment datasets grouped into 3 classes -- positive, neutral, and negative.\r\n\r\nMost multilingual sentiment datasets are either 2-class positive or negative, 5-class ratings of product reviews (e.g. Amazon multilingual dataset), or multiple classes of emotions. However, to an average person, sometimes positive, negative, and neutral classes suffice and are more straightforward to perceive and annotate. Also, a positive/negative classification is too naive, most of the text in the world is neutral in sentiment. Furthermore, most multilingual sentiment datasets don't include Asian languages (e.g. Malay, Indonesian) and are dominated by Western languages (e.g. English, German).\r\n\r\nFor emotions-related datasets, I group the negative (respectively positive) emotions into the negative (respectively positive) class. For rating datasets I assign 1-star reviews to the negative class, 3-star reviews to the neutral class, and assign 5-star reviews to the positive class.","description_withheld":null,"homepage":"https://github.com/tyqiangz/multilingual-sentiment-datasets","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["Multilingual Sentiment Datasets"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}