{"url":"/dataset/nc-sentnob","name":"NC-SentNoB","full_name":"Noise Classification on SentNoB","description_markdown":"This is a multilabel dataset used for Noise Identification purpose in the paper **\"A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts\"** accepted in *2024 The 9th Workshop on Noisy and User-generated Text (W-NUT) collocated with EACL 2024*.\r\n\r\n- Annotated by 4 native Bangla speakers with 90% trustworthiness score.\r\n- Fleiss' Kappa Score: 0.69\r\n\r\n## Definition of noise categories\r\n|Type|Definition|\r\n|-----|---------|\r\n|**Local Word**|Any regional words even if there is a spelling error.|\r\n|**Word Misuse**|Wrong use of words or unnecessary repetitions of words.|\r\n|**Context/Word Missing**|Not enough information or missing words.|\r\n|**Wrong Serial**|Wrong order of the words.|\r\n|**Mixed Language**|Words in another language. Foreign words that were adopted into the Bangla language over time are excluded from this type.|\r\n|**Punctuation Error**|Improper placement or missing punctuation. Sentences ending without \"।\" were excluded from this type.|\r\n|**Spacing Error**|Improper use of white space.|\r\n|**Spelling Error**|Words not following spelling of Bangla Academy Dictionary.|\r\n|**Coined Word**|Emoji, symbolic emoji, link.|\r\n|**Others**|Noises that do not fall into categories mentioned above.|\r\n\r\n\r\n## Statistics of NC-SentNoB per noise class\r\n|Class|Instances|#Word/Instance|\r\n|---|---|---|\r\n|**Local Word**|2,084 (0.136%)|16.05|\r\n|**Word Misuse**|661 (0.043%)|18.55|\r\n|**Context/Word Missing**|550 (0.036%)|13.19|\r\n|**Wrong Serial**|69 (0.005%)|15.30\r\n|**Mixed Language**|6,267 (0.410%)|17.91\r\n|**Punctuation Error**|5,988 (0.391%)|17.25|\r\n|**Spacing Error**|2,456 (0.161%)|18.78|\r\n|**Spelling Error**|5,817 (0.380%)|17.30|\r\n|**Coined Words**|549 (0.036%|15.45|\r\n|**Others**|1,263 (0.083%)|16.52|\r\n\r\n## Heatmap of correlation coefficient\r\n<img src=\"https://huggingface.co/datasets/ktoufiquee/NC-SentNoB/resolve/main/corr_heatmap.png\">\r\n\r\n## Citation\r\nIf you use the datasets, please cite the following paper:\r\n```\r\n@misc{elahi2024comparative,\r\n      title={A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts}, \r\n      author={Kazi Toufique Elahi and Tasnuva Binte Rahman and Shakil Shahriar and Samir Sarker and Md. Tanvir Rouf Shawon and G. M. Shahariar},\r\n      year={2024},\r\n      eprint={2401.14360},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL}\r\n}\r\n```","description_withheld":null,"homepage":"https://github.com/ktoufiquee/A-Comparative-Analysis-of-Noise-Reduction-Methods-in-Sentiment-Analysis-on-Noisy-Bangla-Texts","introduced_date":"2024-01-25","introduced_date_note":null,"introduced_by":{"paper":"/paper/a-comparative-analysis-of-noise-reduction","title":"A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts","first_author":"Kazi Toufique Elahi","url":null},"license":{"name":"cc-by-sa-4.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"Bengali","url":"/datasets/language/bengali"}],"variants":["NC-SentNoB"],"data_loaders":[{"repo":"https://github.com/ktoufiquee/a-comparative-analysis-of-noise-reduction-methods-in-sentiment-analysis-on-noisy-bangla-texts","url":"https://huggingface.co/datasets/ktoufiquee/NC-SentNoB/","frameworks":[]},{"repo":"https://github.com/ktoufiquee/a-comparative-analysis-of-noise-reduction-methods-in-sentiment-analysis-on-noisy-bangla-texts","url":"https://github.com/ktoufiquee/a-comparative-analysis-of-noise-reduction-methods-in-sentiment-analysis-on-noisy-bangla-texts","frameworks":[]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}