{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/building-a-sentiment-corpus-of-tweets-in-1","title":"Building a Sentiment Corpus of Tweets in Brazilian Portuguese","arxiv_id":"1712.08917","date":"2017-12-24","proceeding":null,"authors":["Henrico Bertini Brum","Maria das Graças Volpe Nunes"],"abstract":"The large amount of data available in social media, forums and websites\nmotivates researches in several areas of Natural Language Processing, such as\nsentiment analysis. The popularity of the area due to its subjective and\nsemantic characteristics motivates research on novel methods and approaches for\nclassification. Hence, there is a high demand for datasets on different domains\nand different languages. This paper introduces TweetSentBR, a sentiment corpora\nfor Brazilian Portuguese manually annotated with 15.000 sentences on TV show\ndomain. The sentences were labeled in three classes (positive, neutral and\nnegative) by seven annotators, following literature guidelines for ensuring\nreliability on the annotation. We also ran baseline experiments on polarity\nclassification using three machine learning methods, reaching 80.99% on\nF-Measure and 82.06% on accuracy in binary classification, and 59.85% F-Measure\nand 64.62% on accuracy on three point classification.","url_abs":"http://arxiv.org/abs/1712.08917v1","url_pdf":"http://arxiv.org/pdf/1712.08917v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"building-a-sentiment-corpus-of-tweets-in-1","repo_url":"https://bitbucket.org/HBrum/tweetsentbr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[{"slug":"tweetsentbr","name":"tweetSentBR","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1712.08917","atlas_url":"https://app.syntology.ai/?focus=1712.08917","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}