{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sentnob-a-dataset-for-analysing-sentiment-on","title":"SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts","arxiv_id":null,"date":"2021-11-01","proceeding":"Findings (EMNLP) 2021 11","authors":["Khondoker Ittehadul Islam","Sudipta Kar","Md Saiful Islam","Mohammad Ruhul Amin"],"abstract":"In this paper, we propose an annotated sentiment analysis dataset made of informally written Bangla texts. This dataset comprises public comments on news and videos collected from social media covering 13 different domains, including politics, education, and agriculture. These comments are labeled with one of the polarity labels, namely positive, negative, and neutral. One significant characteristic of the dataset is that each of the comments is noisy in terms of the mix of dialects and grammatical incorrectness. Our experiments to develop a benchmark classification system show that hand-crafted lexical features provide superior performance than neural network and pretrained language models. We have made the dataset and accompanying models presented in this paper publicly available at https://git.io/JuuNB.","url_abs":"https://aclanthology.org/2021.findings-emnlp.278","url_pdf":"https://aclanthology.org/2021.findings-emnlp.278.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sentnob-a-dataset-for-analysing-sentiment-on","repo_url":"https://github.com/KhondokerIslam/SentNoB","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[{"slug":"sentnob","name":"SentNoB","full_name":"SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}