{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-large-scale-dataset-for-hate-speech","title":"A Large-scale Dataset for Hate Speech Detection on Vietnamese Social Media Texts","arxiv_id":"2103.11528","date":"2021-03-22","proceeding":null,"authors":["Son T. Luu","Kiet Van Nguyen","Ngan Luu-Thuy Nguyen"],"abstract":"In recent years, Vietnam witnesses the mass development of social network users on different social platforms such as Facebook, Youtube, Instagram, and Tiktok. On social medias, hate speech has become a critical problem for social network users. To solve this problem, we introduce the ViHSD - a human-annotated dataset for automatically detecting hate speech on the social network. This dataset contains over 30,000 comments, each comment in the dataset has one of three labels: CLEAN, OFFENSIVE, or HATE. Besides, we introduce the data creation process for annotating and evaluating the quality of the dataset. Finally, we evaluated the dataset by deep learning models and transformer models.","url_abs":"https://arxiv.org/abs/2103.11528v4","url_pdf":"https://arxiv.org/pdf/2103.11528v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-large-scale-dataset-for-hate-speech","repo_url":"https://github.com/sonlam1102/vihsd-vietnamese-hate-speech-detection-dataset","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"a-large-scale-dataset-for-hate-speech","repo_url":"https://github.com/sonlam1102/vihsd","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"hate-speech-detection","task_name":"Hate Speech Detection"},{"task_slug":"vietnamese-hate-speech-detection","task_name":"Vietnamese Hate Speech Detection"},{"task_slug":"vietnamese-social-media-text-processing","task_name":"Vietnamese Social Media Text Processing"}],"methods":[],"datasets_introduced":[{"slug":"vihsd","name":"ViHSD","full_name":"Vietnamese Hate Speech Detection Dataset"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2103.11528","atlas_url":"https://app.syntology.ai/?focus=2103.11528","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}