{"url":"/dataset/tweetsentbr","name":"tweetSentBR","full_name":null,"description_markdown":"The **TweetSentBR Dataset** is a valuable resource for sentiment analysis in Brazilian Portuguese. Let me provide you with some details about it:\r\n\r\n1. **Description**:\r\n   - The dataset consists of **15,000 manually annotated sentences** extracted from tweets in Brazilian Portuguese.\r\n   - These sentences are specifically related to the **TV show domain**.\r\n   - Each sentence has been labeled into one of three classes: **positive**, **neutral**, or **negative** sentiment.\r\n   - The annotation process followed literature guidelines to ensure reliability.\r\n\r\n2. **Purpose**:\r\n   - Researchers and practitioners in the field of **Natural Language Processing (NLP)** use this dataset for sentiment analysis tasks.\r\n   - It serves as a benchmark for developing and evaluating novel methods and approaches for sentiment classification.\r\n\r\n3. **Performance**:\r\n   - Baseline experiments on polarity classification using three machine learning methods achieved the following results:\r\n     - **Binary classification** (positive vs. negative): **80.99% F-Measure** and **82.06% accuracy**.\r\n     - **Three-point classification** (positive, neutral, negative): **59.85% F-Measure** and **64.62% accuracy**.\r\n\r\nSource: Conversation with Bing, 3/16/2024\r\n(1) Building a Sentiment Corpus of Tweets in Brazilian Portuguese. https://arxiv.org/abs/1712.08917.\r\n(2) 7 Best Portuguese Language Speech Datasets of 2022 | Twine. https://www.twine.net/blog/portuguese-language-speech-datasets/.\r\n(3) A survey and study impact of tweet sentiment analysis via ... - Springer. https://link.springer.com/article/10.1007/s10579-023-09687-8.\r\n(4) Top 25 Twitter Datasets for NLP and Machine Learning | iMerit. https://imerit.net/blog/top-25-twitter-datasets-for-natural-language-processing-and-machine-learning-all-pbm/.\r\n(5) Building a Sentiment Corpus of Tweets in Brazilian Portuguese - arXiv.org. https://arxiv.org/pdf/1712.08917v1.pdf.\r\n(6) undefined. https://doi.org/10.48550/arXiv.1712.08917.","description_withheld":null,"homepage":"","introduced_date":"2017-12-24","introduced_date_note":null,"introduced_by":{"paper":"/paper/building-a-sentiment-corpus-of-tweets-in-1","title":"Building a Sentiment Corpus of Tweets in Brazilian Portuguese","first_author":"Henrico Bertini Brum","url":null},"license":null,"modalities":[],"tasks":[{"name":"Text Generation","url":"/task/text-generation","datasets_with_task":"/datasets/task/text-generation"}],"languages":[],"variants":["tweetSentBR"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-generation-on-tweetsentbr","task":"Text Generation","dataset_variant":"tweetSentBR","rows":0,"metrics":["f1-macro"],"first_row_in_archive_order":null,"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}